Beginner 8 minConcepts

GGUF, safetensors: understanding the formats of models

You open a model page on Hugging Face and find a flood of files: .safetensors, sometimes .gguf, names like Q4_K_M or model-00001-of-00004. Which file should you download? This guide decodes the two formats that really matter today — GGUF for local inference, safetensors on Hugging Face — explains why Ollama and LM Studio require GGUF, how to read a filename correctly, and how to convert from one format to the other when necessary.

By Samir K.·Update 2026-09-20·Tested on Windows, macOS, and Linux

#Why these formats exist

Once trained, an LLM is nothing more than a huge bag of numbers: the weights, meaning the billions of parameters that encode what the model “knows.” The file format is simply how those numbers are arranged on disk. You have to decide how to store them, how to read them quickly, and which additional information—the vocabulary, architecture, and settings—to include alongside them.

Historically, these weights were saved in PyTorch’s .bin format (Python pickle), which was convenient but slow to load and, above all, dangerous: a pickle file can execute arbitrary code when opened. Two formats emerged to solve different problems. safetensors meets the needs of researchers and platforms: safe, fast storage that preserves the original precision. GGUF meets the needs of local inference: a single, compact, quantized file ready to run on a mainstream CPU or GPU.

i
Format ≠ model
The same model (say Qwen 3.5 7B) exists in several formats. It isn't a different model: they're the same weights, arranged differently. Choosing GGUF over safetensors doesn't change the model's intelligence, only how it runs.

#safetensors: the Hugging Face format

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

safetensors is the default format in the Hugging Face ecosystem. Created to replace PyTorch pickle, its primary value is right there in its name: safety. The file contains only data (the weight tensors) and a small JSON header describing their shape and type. No executable code, so opening a malicious file poses no risk. Bonus: loading is faster thanks to memory mapping, which lets you read the weights directly from disk without copying everything into RAM first.

Safe by design
No executable code, unlike older .bin/.pt files using pickle. You can download a safetensors file without worrying about code execution.
Full precision
Weights are generally stored in FP16 or BF16 (16-bit), or even FP32. This is the training precision, the reference standard.
Built to be transformed
This is the starting format for fine-tuning, model merging, quantization, or conversion to other formats.
Often split
A large model is split into multiple files (shards) accompanied by a JSON index because a single file of tens of GB would be unmanageable.

The downside: an FP16 safetensors file is large. A 7B model weighs about 14 GB (2 bytes per parameter), while a 70B model weighs around 140 GB. For training and research on server GPUs, it’s the right choice. To run a model on your machine, it’s often too heavy—which is why quantization and GGUF are useful.

→
The rule of thumb for size
In FP16, expect about 2 GB of file size per billion parameters. A 7B model is ≈ 14 GB, and a 13B model is ≈ 26 GB. After Q4 quantization (GGUF), that drops to around 0.6–0.7 GB per billion: the 7B drops to ~4–5 GB.

#GGUF: the format for local inference

GGUF (GPT-Generated Unified Format) is the format born from the llama.cpp project, the inference engine that runs LLMs efficiently on both CPUs and GPUs. It replaced the older GGML format in 2023. Its central idea: put everything in a single file. The weights, tokenizer vocabulary, architecture metadata (number of layers, context size), and chat template—all packaged together. You download one file, run it, and it works.

Single, self-contained file
Weights, tokenizer, and metadata in a single .gguf file. No config directory to assemble and no Python dependency to install.
Quantized
Designed to accommodate compressed weights (4, 5, 6, and 8 bits). This is what makes large models accessible on consumer hardware.
CPU + GPU + offload
llama.cpp can split layers between the GPU and system RAM. You can run a model larger than your VRAM, at the cost of some speed.
Portable
The same .gguf file works on Windows, macOS (Metal), and Linux, with Ollama, LM Studio, Jan, or llama.cpp directly.

Quantization is the heart of the matter. It stores each weight using fewer bits—4 instead of 16, for example—which cuts the size by a factor of 3 to 4 with minimal quality loss when chosen well. This is what lets a 7B model fit in ~5 GB of VRAM instead of 14. Common levels include: Q4_K_M (the best compromise, recommended by default), Q5_K_M (a step up in quality), Q8_0 (almost lossless, but heavier), and FP16 (unquantized, the reference).

i
GGUF is not necessarily quantized
You can also produce a GGUF in FP16, without quantization. But in 99% of cases, if you download a GGUF, it's a quantized version: that's precisely what the format excels at.

#Why Ollama and LM Studio need GGUF

Ollama and LM Studio are built on top of llama.cpp (or an equivalent engine), and llama.cpp natively supports GGUF. This isn't a gimmick: it's what makes these tools so simple. Since GGUF already contains the tokenizer, architecture, and chat template, the tool has nothing to guess. It reads the file, allocates memory, and responds. No Python environment, no dependencies to resolve, no configuration to write.

When you run `ollama pull llama3.2`, Ollama actually downloads a GGUF from its registry and stores it in its model repository. You never see the file, but it is indeed GGUF under the hood. LM Studio, on the other hand, explicitly shows you the available GGUF files and their quantizations at download time.

Terminal
# Ollama récupère un GGUF depuis sa registry et le sert sur localhost:11434
ollama pull llama3.2
ollama run llama3.2

# Vérifier que le daemon répond
curl http://localhost:11434/api/tags
→
Load an external GGUF into Ollama
Did you manually download a .gguf (from Hugging Face, for example)? Ollama can import it through a two-line Modelfile: `FROM ./mon-modele.gguf`, then `ollama create mon-modele -f Modelfile`. LM Studio, on the other hand, automatically detects .gguf files placed in its model folder.

#Read a filename correctly

On Hugging Face, GGUF filenames follow a readable convention once you know the code. Take a typical example: `Qwen2.5-7B-Instruct-Q4_K_M.gguf`. Each part carries information.

Qwen2.5
The model family and version.
7B
The parameter count: 7 billion. This is the first indicator of the required VRAM.
Instruct
The variant trained to follow instructions and hold conversations (as opposed to -base, raw, and not aligned for chat).
Q4_K_M
Quantization: 4-bit, K_M variant (medium). The default-recommended quality/size tradeoff.
.gguf
The format. You know it will run in Ollama, LM Studio, or llama.cpp without conversion.

The quantization suffix is the most useful part to decode. The number indicates the number of bits per weight; the letters K_S / K_M / K_L designate variants (Small, Medium, Large) that protect sensitive layers to varying degrees. The higher the number, the larger and more faithful the file.

Q4_K_M
~4 bits, medium. The default choice: the best quality for the size in the vast majority of cases.
Q5_K_M
~5 bits. One step up in quality, with a slightly larger file. Good if VRAM allows it.
Q8_0
8 bits. Nearly indistinguishable from non-quantized, but twice as large as Q4. For purists or demanding tasks.
Q2_K / Q3_K
2–3 bits. Very compact but with visibly degraded quality. Reserve it for cases where memory is truly critical.
FP16 / F16
Unquantized, full 16-bit precision. The reference, but heavy—so you might as well stick with safetensors in this case.
!
The trap of a model split into shards
An oversized GGUF is sometimes split: `model-00001-of-00002.gguf`, `model-00002-of-00002.gguf`. You must download ALL the pieces and keep them in the same folder—the engine reassembles them. Do not take a single file from a split series thinking you have the complete model.

#Which one to download for your tool

The practical question boils down to: what tool will I use? The format follows from the answer, not the other way around.

Ollama, LM Studio, Jan, llama.cpp
→ GGUF. These tools are built for this. Choose the quantization according to your VRAM (Q4_K_M by default).
vLLM, TGI, Transformers (Python)
→ safetensors. These server engines load the native Hugging Face format, often in FP16 or with their own quantization (AWQ, GPTQ).
Fine-tuning, merging, custom quantization
→ safetensors. This is the working format: start from full precision to transform the model.
You don't know yet
→ Quantized GGUF if you’re using it locally on your machine. It’s the simplest and most memory-efficient option.

To choose the right quantization, match it to your VRAM. Q4 guidelines: a 3B model fits in ~2 GB, a 7B in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B in ~40 GB. A RTX 3060 12 GB comfortably runs 7B to 14B models in Q4; a RTX 4090 24 GB targets 32B; for 70B, aim for a card with large VRAM or a Mac with unified memory (M4 Pro 24-48 GB).

→
When in doubt, Q4_K_M
Nine times out of ten, the Q4_K_M version is the right file: it offers the best quality-to-memory ratio and fits on most consumer GPUs. Move up to Q5_K_M or Q8_0 only if you have VRAM to spare and a specific quality requirement.

#Convert safetensors to GGUF

Sometimes a model is published only in safetensors (common on release day), and you want to run it in Ollama. You then need to convert it to GGUF and possibly quantize it. The reference tool is the `convert_hf_to_gguf.py` script provided by llama.cpp. The process takes two steps: first convert to full-precision GGUF, then quantize with the `llama-quantize` tool.

  1. 01
    Fetch llama.cpp and its dependencies
    Clone the llama.cpp repository and install the Python dependencies for the conversion script. That script contains convert_hf_to_gguf.py.
  2. 02
    Download the safetensors model
    Download the complete model folder from Hugging Face (.safetensors weights + config.json + tokenizer files). Everything must be present, not just the weights.
  3. 03
    Convert to GGUF FP16
    Run convert_hf_to_gguf.py on the model directory. You get a full-precision (16-bit) .gguf file that is large but faithful.
  4. 04
    Quantize in Q4_K_M
    Run the GGUF FP16 through llama-quantize and choose the desired level (Q4_K_M by default). The final file is 3 to 4 times smaller.
  5. 05
    Import into Ollama
    Write a Modelfile pointing to the quantized .gguf and create the model with ollama create. It can then be used like any other Ollama model.
Terminal
# 1. Cloner llama.cpp et installer les dépendances de conversion
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
pip install -r requirements.txt

# 2. Convertir le dossier safetensors en GGUF pleine précision
python convert_hf_to_gguf.py ./mon-modele-hf --outfile mon-modele-f16.gguf --outtype f16

# 3. Quantifier en Q4_K_M (compromis recommandé)
./llama-quantize mon-modele-f16.gguf mon-modele-Q4_K_M.gguf Q4_K_M
Modelfile + import Ollama
# Créer un Modelfile minimal
printf 'FROM ./mon-modele-Q4_K_M.gguf\n' > Modelfile

# Enregistrer le modèle dans Ollama
ollama create mon-modele -f Modelfile
ollama run mon-modele
!
Conversion doesn't create quality
Converting and then quantizing does not make a model better—in fact, each quantization step removes a bit of precision. If an official GGUF version already exists (often published by the community, such as the “GGUF” repositories on Hugging Face), download it instead of converting it yourself: it’s faster and often better calibrated.

#The other formats you’ll come across

GGUF and safetensors cover most use cases, but a few other names appear during downloads. Knowing them helps avoid unpleasant surprises.

.bin / .pt (pickle)
The old PyTorch format. Functional but unsafe (it can execute code). Gradually replaced by safetensors — avoid it when an alternative exists.
GPTQ / AWQ
GPU-side quantizations for vLLM and Transformers, stored in safetensors. Fast on GPU NVIDIA, but not readable by Ollama/llama.cpp.
MLX
The Apple format for its MLX framework, optimized for Apple Silicon. Used by some native Mac apps, distinct from GGUF.
ONNX
A multi-framework exchange format, mainly for industrial and edge deployment. Rare for mainstream local LLM use.
GGML
The ancestor of GGUF (same project). Obsolete: if you encounter a .ggml file, look for the equivalent .gguf version.

#Frequently asked questions

GGUF or safetensors—which is better?
Neither in absolute terms: they serve different use cases. GGUF for running a model locally (Ollama, LM Studio), safetensors for the Hugging Face ecosystem, fine-tuning, and server engines. The “best” option depends on your tool.
Is a GGUF worse than safetensors?
A quantized GGUF loses a little precision compared with the original FP16 safetensors file. In Q4_K_M or Q5_K_M, the difference is minor and rarely noticeable in practice. In Q2/Q3, it becomes visible.
Can I use a safetensors file directly in Ollama?
Not directly in most cases: Ollama expects GGUF. You must first convert safetensors to GGUF with llama.cpp. Some recent versions accept safetensors imports, but GGUF remains the reliable route.
Why are there so many files on a Hugging Face page?
A model is often split into multiple shards (safetensors or GGUF), plus config and tokenizer files. For GGUF, you generally want a single file per quantization (or all the pieces of a split series).

#Go further

Now that the formats hold no secrets, these guides naturally take the topic further.

Choose your quantization (Q4, Q5, Q8, FP16)
The detailed guide for choosing between GGUF quantization levels based on your VRAM and quality needs.
What is Ollama and how does it work
Understand the tool that downloads and serves GGUFs locally, with the basic commands.
Understanding the context window
The other parameter that affects memory is tokens and context, combined with the quantization choice.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.