GGUF, safetensors: understanding the formats of models
You open a model page on Hugging Face and find a flood of files: .safetensors, sometimes .gguf, names like Q4_K_M or model-00001-of-00004. Which file should you download? This guide decodes the two formats that really matter today — GGUF for local inference, safetensors on Hugging Face — explains why Ollama and LM Studio require GGUF, how to read a filename correctly, and how to convert from one format to the other when necessary.
#Why these formats exist
Once trained, an LLM is nothing more than a huge bag of numbers: the weights, meaning the billions of parameters that encode what the model “knows.” The file format is simply how those numbers are arranged on disk. You have to decide how to store them, how to read them quickly, and which additional information—the vocabulary, architecture, and settings—to include alongside them.
Historically, these weights were saved in PyTorch’s .bin format (Python pickle), which was convenient but slow to load and, above all, dangerous: a pickle file can execute arbitrary code when opened. Two formats emerged to solve different problems. safetensors meets the needs of researchers and platforms: safe, fast storage that preserves the original precision. GGUF meets the needs of local inference: a single, compact, quantized file ready to run on a mainstream CPU or GPU.
#safetensors: the Hugging Face format
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
safetensors is the default format in the Hugging Face ecosystem. Created to replace PyTorch pickle, its primary value is right there in its name: safety. The file contains only data (the weight tensors) and a small JSON header describing their shape and type. No executable code, so opening a malicious file poses no risk. Bonus: loading is faster thanks to memory mapping, which lets you read the weights directly from disk without copying everything into RAM first.
- Safe by design
- No executable code, unlike older .bin/.pt files using pickle. You can download a safetensors file without worrying about code execution.
- Full precision
- Weights are generally stored in FP16 or BF16 (16-bit), or even FP32. This is the training precision, the reference standard.
- Built to be transformed
- This is the starting format for fine-tuning, model merging, quantization, or conversion to other formats.
- Often split
- A large model is split into multiple files (shards) accompanied by a JSON index because a single file of tens of GB would be unmanageable.
The downside: an FP16 safetensors file is large. A 7B model weighs about 14 GB (2 bytes per parameter), while a 70B model weighs around 140 GB. For training and research on server GPUs, it’s the right choice. To run a model on your machine, it’s often too heavy—which is why quantization and GGUF are useful.
#GGUF: the format for local inference
GGUF (GPT-Generated Unified Format) is the format born from the llama.cpp project, the inference engine that runs LLMs efficiently on both CPUs and GPUs. It replaced the older GGML format in 2023. Its central idea: put everything in a single file. The weights, tokenizer vocabulary, architecture metadata (number of layers, context size), and chat template—all packaged together. You download one file, run it, and it works.
- Single, self-contained file
- Weights, tokenizer, and metadata in a single .gguf file. No config directory to assemble and no Python dependency to install.
- Quantized
- Designed to accommodate compressed weights (4, 5, 6, and 8 bits). This is what makes large models accessible on consumer hardware.
- CPU + GPU + offload
- llama.cpp can split layers between the GPU and system RAM. You can run a model larger than your VRAM, at the cost of some speed.
- Portable
- The same .gguf file works on Windows, macOS (Metal), and Linux, with Ollama, LM Studio, Jan, or llama.cpp directly.
Quantization is the heart of the matter. It stores each weight using fewer bits—4 instead of 16, for example—which cuts the size by a factor of 3 to 4 with minimal quality loss when chosen well. This is what lets a 7B model fit in ~5 GB of VRAM instead of 14. Common levels include: Q4_K_M (the best compromise, recommended by default), Q5_K_M (a step up in quality), Q8_0 (almost lossless, but heavier), and FP16 (unquantized, the reference).
#Why Ollama and LM Studio need GGUF
Ollama and LM Studio are built on top of llama.cpp (or an equivalent engine), and llama.cpp natively supports GGUF. This isn't a gimmick: it's what makes these tools so simple. Since GGUF already contains the tokenizer, architecture, and chat template, the tool has nothing to guess. It reads the file, allocates memory, and responds. No Python environment, no dependencies to resolve, no configuration to write.
When you run `ollama pull llama3.2`, Ollama actually downloads a GGUF from its registry and stores it in its model repository. You never see the file, but it is indeed GGUF under the hood. LM Studio, on the other hand, explicitly shows you the available GGUF files and their quantizations at download time.
#Read a filename correctly
On Hugging Face, GGUF filenames follow a readable convention once you know the code. Take a typical example: `Qwen2.5-7B-Instruct-Q4_K_M.gguf`. Each part carries information.
- Qwen2.5
- The model family and version.
- 7B
- The parameter count: 7 billion. This is the first indicator of the required VRAM.
- Instruct
- The variant trained to follow instructions and hold conversations (as opposed to -base, raw, and not aligned for chat).
- Q4_K_M
- Quantization: 4-bit, K_M variant (medium). The default-recommended quality/size tradeoff.
- .gguf
- The format. You know it will run in Ollama, LM Studio, or llama.cpp without conversion.
The quantization suffix is the most useful part to decode. The number indicates the number of bits per weight; the letters K_S / K_M / K_L designate variants (Small, Medium, Large) that protect sensitive layers to varying degrees. The higher the number, the larger and more faithful the file.
- Q4_K_M
- ~4 bits, medium. The default choice: the best quality for the size in the vast majority of cases.
- Q5_K_M
- ~5 bits. One step up in quality, with a slightly larger file. Good if VRAM allows it.
- Q8_0
- 8 bits. Nearly indistinguishable from non-quantized, but twice as large as Q4. For purists or demanding tasks.
- Q2_K / Q3_K
- 2–3 bits. Very compact but with visibly degraded quality. Reserve it for cases where memory is truly critical.
- FP16 / F16
- Unquantized, full 16-bit precision. The reference, but heavy—so you might as well stick with safetensors in this case.
#Which one to download for your tool
The practical question boils down to: what tool will I use? The format follows from the answer, not the other way around.
- Ollama, LM Studio, Jan, llama.cpp
- → GGUF. These tools are built for this. Choose the quantization according to your VRAM (Q4_K_M by default).
- vLLM, TGI, Transformers (Python)
- → safetensors. These server engines load the native Hugging Face format, often in FP16 or with their own quantization (AWQ, GPTQ).
- Fine-tuning, merging, custom quantization
- → safetensors. This is the working format: start from full precision to transform the model.
- You don't know yet
- → Quantized GGUF if you’re using it locally on your machine. It’s the simplest and most memory-efficient option.
To choose the right quantization, match it to your VRAM. Q4 guidelines: a 3B model fits in ~2 GB, a 7B in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B in ~40 GB. A RTX 3060 12 GB comfortably runs 7B to 14B models in Q4; a RTX 4090 24 GB targets 32B; for 70B, aim for a card with large VRAM or a Mac with unified memory (M4 Pro 24-48 GB).
#Convert safetensors to GGUF
Sometimes a model is published only in safetensors (common on release day), and you want to run it in Ollama. You then need to convert it to GGUF and possibly quantize it. The reference tool is the `convert_hf_to_gguf.py` script provided by llama.cpp. The process takes two steps: first convert to full-precision GGUF, then quantize with the `llama-quantize` tool.
- 01Fetch llama.cpp and its dependenciesClone the llama.cpp repository and install the Python dependencies for the conversion script. That script contains convert_hf_to_gguf.py.
- 02Download the safetensors modelDownload the complete model folder from Hugging Face (.safetensors weights + config.json + tokenizer files). Everything must be present, not just the weights.
- 03Convert to GGUF FP16Run convert_hf_to_gguf.py on the model directory. You get a full-precision (16-bit) .gguf file that is large but faithful.
- 04Quantize in Q4_K_MRun the GGUF FP16 through llama-quantize and choose the desired level (Q4_K_M by default). The final file is 3 to 4 times smaller.
- 05Import into OllamaWrite a Modelfile pointing to the quantized .gguf and create the model with ollama create. It can then be used like any other Ollama model.
#The other formats you’ll come across
GGUF and safetensors cover most use cases, but a few other names appear during downloads. Knowing them helps avoid unpleasant surprises.
- .bin / .pt (pickle)
- The old PyTorch format. Functional but unsafe (it can execute code). Gradually replaced by safetensors — avoid it when an alternative exists.
- GPTQ / AWQ
- GPU-side quantizations for vLLM and Transformers, stored in safetensors. Fast on GPU NVIDIA, but not readable by Ollama/llama.cpp.
- MLX
- The Apple format for its MLX framework, optimized for Apple Silicon. Used by some native Mac apps, distinct from GGUF.
- ONNX
- A multi-framework exchange format, mainly for industrial and edge deployment. Rare for mainstream local LLM use.
- GGML
- The ancestor of GGUF (same project). Obsolete: if you encounter a .ggml file, look for the equivalent .gguf version.
#Frequently asked questions
- GGUF or safetensors—which is better?
- Neither in absolute terms: they serve different use cases. GGUF for running a model locally (Ollama, LM Studio), safetensors for the Hugging Face ecosystem, fine-tuning, and server engines. The “best” option depends on your tool.
- Is a GGUF worse than safetensors?
- A quantized GGUF loses a little precision compared with the original FP16 safetensors file. In Q4_K_M or Q5_K_M, the difference is minor and rarely noticeable in practice. In Q2/Q3, it becomes visible.
- Can I use a safetensors file directly in Ollama?
- Not directly in most cases: Ollama expects GGUF. You must first convert safetensors to GGUF with llama.cpp. Some recent versions accept safetensors imports, but GGUF remains the reliable route.
- Why are there so many files on a Hugging Face page?
- A model is often split into multiple shards (safetensors or GGUF), plus config and tokenizer files. For GGUF, you generally want a single file per quantization (or all the pieces of a split series).
#Go further
Now that the formats hold no secrets, these guides naturally take the topic further.
- Choose your quantization (Q4, Q5, Q8, FP16)
- The detailed guide for choosing between GGUF quantization levels based on your VRAM and quality needs.
- What is Ollama and how does it work
- Understand the tool that downloads and serves GGUFs locally, with the basic commands.
- Understanding the context window
- The other parameter that affects memory is tokens and context, combined with the quantization choice.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.