Import a GGUF model from Hugging Face into Ollama
The official Ollama library covers only a fraction of the available models. On Hugging Face, tens of thousands of GGUF files are waiting—community fine-tunes, recent models, and versions not yet packaged. This guide shows how to import any Hugging Face GGUF into Ollama: the direct ollama run hf.co command, the Modelfile FROM method for a local file, how to choose quantization based on your VRAM, and how to fix a broken chat template that makes responses inconsistent.
#Why import a GGUF from Hugging Face
Ollama maintains a library of practical models (ollama.com/library), but it is intentionally limited: the maintainers publish the most requested models there, using quantizations selected for them. As soon as you are looking for a specialized fine-tune, a newly released version, a French-language model, or a specific quantization, you have to find it on Hugging Face—the largest platform for sharing open-weight models.
The GGUF format (successor to GGML) is the format that Ollama understands natively: a single file containing the quantized weights, tokenizer, and model metadata. Contributors such as TheBloke, bartowski, and unsloth publish thousands of ready-to-use GGUFs, often in around ten quantizations per model. Knowing how to import them unlocks this entire ecosystem in your Ollama installation.
- Recent models
- A model published yesterday on Hugging Face can be used before it even appears in Ollama's official library.
- Niche fine-tunes
- Specialized models (code, medicine, role-playing, French) that no one has taken the trouble to package officially.
- Precise quantization
- Choose exactly the level (Q4_K_M, Q5_K_M, Q8_0…) that fits in your VRAM, rather than settling for the default variant.
- Private models
- Your own fine-tunes or downloaded GGUF files, imported locally through a Modelfile.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Ollama installed
- Recent version (0.5+) for native hf.co support. The daemon listens on http://localhost:11434 by default. Check with “ollama --version”.
- An internet connection
- For the direct method that downloads from Hugging Face. After that, the model runs 100% locally.
- Enough VRAM or RAM
- Q4 rule of thumb: a 7B fits in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B in ~40 GB. Without a GPU, RAM is what matters, and it will be slower.
- The name of a GGUF repository
- For example, bartowski/Qwen3.5-9B-Instruct-GGUF. Find it in the URL of the model's Hugging Face page.
To find a GGUF repository, Hugging Face search supports filtering by format. Search for the model name and add “GGUF,” or filter for the “GGUF” library in the sidebar. Open the “Files and versions” tab: you'll see the list of .gguf files, one per quantization, with their size in GB — valuable information for what comes next.
#Direct method: ollama run hf.co/...
This is by far the simplest way to import a GGUF model from Hugging Face into Ollama. Since version 0.5, Ollama can pull a GGUF directly from a Hugging Face repository with a single command, without downloading the file manually or writing a Modelfile. The syntax uses the repository path prefixed with hf.co/.
Without a specification, Ollama chooses a default quantization (usually Q4_K_M if it exists in the repository). To target a specific quantization, add it after a colon, exactly like a standard model tag. The tag name matches the .gguf filename suffix and is case-insensitive.
Ollama downloads the file, saves it to its local storage, and starts the conversation. The model then appears in « ollama list » under its full hf.co/... name and can be restarted instantly. You can give it a shorter alias with « ollama cp » if the name feels too long to type.
#Choosing the right quantization for your VRAM
The same model is released in several quantizations: this is the central trade-off between quality and memory. The more aggressive the quantization (fewer bits per weight), the smaller the file and the more easily it fits on a modest card—at the cost of a slight loss in precision. The right approach is to choose the highest quantization that fits comfortably in your VRAM.
- Q4_K_M — recommended
- The best trade-off for the vast majority of use cases. Almost imperceptible quality loss, with a smaller memory footprint. Choose it by default if you're unsure.
- Q5_K_M — one step up
- Slightly heavier, slightly more accurate. Interesting if you have VRAM to spare and want maximum quality without moving to 8-bit.
- Q8_0 — nearly lossless
- Very close to the unquantized model, but about twice as large as Q4. Reserved for cases where even the slightest degradation matters and VRAM is plentiful.
- FP16 — full precision
- The unquantized model, and the heaviest one. Rarely necessary for local inference: Q8_0 is almost always sufficient and cuts memory usage in half.
To estimate whether a quantization will fit, use the .gguf file size shown on Hugging Face, plus a margin of about 1 to 2 GB for the context and system. Here are the Q4_K_M VRAM guidelines by model size, along with the typical GPUs that can run them.
- 3B ≈ 2 GB
- Runs everywhere, even on an entry-level card or on the CPU. Ideal RTX 3060 12 GB with context to spare.
- 7B ≈ 5 GB
- Comfortable on RTX 3060 12 GB, RTX 4070 12 GB. The most versatile format for everyday use.
- 14B ≈ 9 GB
- RTX 4070 12 GB (barely), RTX 4080 16 GB comfortably. A good quality tier for reasoning and coding.
- 32B ≈ 19 GB
- RTX 4090 24 GB, or a Mac M4 Pro with unified memory. The accessible high end for a workstation.
- 70B ≈ 40 GB
- Requires 48 GB+: a Mac Studio with large unified memory or a multi-GPU setup. Use Q4 or an even more aggressive quantization.
#Modelfile method: FROM filename.gguf
The direct method assumes that the GGUF is on Hugging Face and accessible online. But if you have already downloaded a .gguf file manually, built your own with llama.cpp, or want to customize the model (system prompt, parameters), you need to use a Modelfile. It is a small text file, similar to a Dockerfile, that describes how to build a Ollama model from a local GGUF.
The core directive is FROM, which points to the .gguf file path. Create a file named “Modelfile” (with no extension) next to your GGUF, containing at least this line.
Then build the model with « ollama create », giving it any name you like. Ollama reads the GGUF, saves it to its storage, and makes it available like any other model.
A complete Modelfile can go much further: set a system prompt, tune sampling parameters, and above all define the TEMPLATE—the chat format expected by the model. This is where most quality issues are fixed, as we'll see in the next section.
#Fix a broken chat template
This is the number-one trap when importing GGUF files. An imported model may respond unpredictably: sentences that never end, strange tags in the output (<|im_end|>, [INST], <end_of_turn>), answers that ignore the question or go into a loop. Nine times out of ten, the model is not bad—the chat template does not match the one used during training.
Each model family expects a specific conversation format: ChatML (<|im_start|>) for Qwen and many fine-tunes, [INST]...[/INST] for Mistral and Llama 2, <start_of_turn> for Gemma, and a specific format for Llama 3. If the GGUF embeds the wrong template in its metadata, or if Ollama guesses an incorrect one, responses degrade. The most common symptom is end-of-turn tags appearing as plain text in the response instead of stopping generation.
- Symptom: visible tags
- The model displays <|im_end|> or <|eot_id|> in its response. A corresponding PARAMETER stop is missing, or the template does not emit the correct end token.
- Symptom: endless generation
- The model never stops and keeps going through turns on its own. The expected stop token is not declared.
- Symptom: inconsistent responses
- The model ignores the system prompt or gives an off-target response. The role format (system/user/assistant) does not match the one used during training.
The fix is to provide the correct TEMPLATE and the correct PARAMETER stop settings in a Modelfile. The source of truth is the original model's Hugging Face “model card”: look for the “prompt format” or “chat template” section, which specifies the exact format. For a ChatML model (Qwen and derivatives), the template and stops look like this.
Then rebuild with “ollama create” and test it. An effective way to retrieve the right template without rewriting it: start from an official model from the same family that is already present in Ollama and inspect its generated Modelfile, then reuse its TEMPLATE block.
#Troubleshooting
- “Error: pull model manifest”
- The hf.co path is misspelled, the repository is private/gated, or your version of Ollama is too old. Check the repository's exact URL and update Ollama.
- The quantization tag does not exist
- Ollama says the tag cannot be found: open “Files and versions” on Hugging Face and copy the exact suffix of the .gguf file (e.g., Q4_K_M, IQ4_XS). Case is ignored, but the name must match.
- Very slow model / choppy responses
- The model spills over into RAM/CPU because it lacks enough VRAM. Check with “ollama ps” whether it’s running on the GPU or CPU, and drop down one quantization level or model size.
- The repository contains only safetensors
- No .gguf available: look for a community-converted “GGUF” version, or convert the model yourself using the llama.cpp scripts.
- Outputs polluted by tags
- Incorrect chat template: see the previous section, and provide the correct TEMPLATE and PARAMETER stop values in a Modelfile.
#Go further
Importing a GGUF relies on two core skills from the Ollama ecosystem. These site guides build on this one:
- Choose your quantization (Q4, Q5, Q8, FP16)
- Understand the quality/memory trade-off in detail so you can choose the right GGUF variant for your GPU.
- Customize a model with the Ollama Modelfile
- Go further with the Modelfile: system prompts, parameters, templates, and multiple variants of the same model.
- Install Ollama: Windows, macOS, and Linux
- The up-to-date basic installation guide, if you're starting from scratch before importing your first GGUFs.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.