Hugging Face: what it is and how to use it ?
Hugging Face is the platform where almost all open AI models are published: Qwen, Llama, Gemma, Mistral, gpt-oss, DeepSeek. It is often described as the GitHub of machine learning: the repository where the files come from, including those that Ollama and LM Studio download for you. The Hub passed three million public models on August 18, 2026, without any editorial review.
Hugging Face is the platform where almost all open AI models are published: Qwen, Llama, Gemma, Mistral, gpt-oss, DeepSeek. It’s often described as the GitHub of machine learning. If you run AI on your own machine, it’s the repository all the files come from, including the files that Ollama and LM Studio download for you. The Hub officially passed the three-million-public-model mark on August 18, 2026, at a rate of about 2,700 to 3,000 new models per day, with no filtering at all. This page teaches you how to read it: which repository to open, which file to choose for your graphics card, and what to check before clicking.
#What is Hugging Face?
Hugging Face is a company founded in 2016 by three French entrepreneurs, Clément Delangue, Julien Chaumond, and Thomas Wolf, and based between New York and Paris, with its headquarters remaining in the United States. It started with a chatbot app for teenagers, a project now abandoned, hence the name borrowed from the emoji that smiles with its hands open. It then turned to open-source tooling for machine learning: its transformers library became the de facto standard way to load and run a language model in Python, adopted by most research labs and companies in the industry. The company eventually built the Hub, a hosting platform based on the Git version-control system where anyone can publish models, datasets, and small web demos for free, without prior editorial review.
For local AI, the Hub acts like an app store, with two key differences: almost everything there is free, and nothing is selected for you. Unlike a traditional app store, nobody tests, approves, or ranks models before publication: anyone can create a free account and upload a repository in a few minutes, without any validation or quality control whatsoever. This complete lack of filtering is both the Hub's strength—everything arrives there first, often even before a lab's official announcement—and its main pitfall for beginners, who may face hundreds of search results without knowing which one is reliable. Knowing how to navigate it, distinguish an official repository from a third-party copy and a serious conversion from an abandoned experiment, is the skill to acquire before downloading anything.
#The Hub’s three sections
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
| Radius | Content | Interest in local AI |
|---|---|---|
| Models | Weight files, configuration, documentation | This is where the file you'll run is located |
| Datasets | Training and evaluation data | Only if you’re doing fine-tuning or testing |
| Spaces | Small web applications, mostly demos | Try a model in the browser before downloading 15 GB |
#Read a model page
Each model lives at an address in the form huggingface.co/organisation/nom-du-modele, for example huggingface.co/Qwen/Qwen3-8B. Four distinct elements on this page deserve your attention before any download.
- The organization name
- Qwen, google, meta-llama, mistralai, openai, deepseek-ai are the labs themselves, verified. Any other name refers to a third party: sometimes excellent, but never the lab's original.
- The model card
- The README file: what the model is, how it was trained, its intended uses, its scores, its prompt format, and its limitations. Quality ranges from highly comprehensive to empty.
- The license
- Displayed prominently at the top of the model page, next to the model name. See the dedicated section below for what each common license actually allows.
- The Files and versions tab
- The complete list of downloadable files. A quick glance is enough to immediately tell what type of repository you are looking at.
#Which repository to choose: original weights or GGUF
| What you see in Files | What it is | For which tools | Size for an 8B model |
|---|---|---|---|
| model-00001-of-00004.safetensors… | Original weights, in BF16, split across multiple files | transformers, vLLM, fine-tuning tools | ≈ 16 GB |
| …-Q4_K_M.gguf, …-Q8_0.gguf | Quantized conversions, one file per level | Ollama, LM Studio, llama.cpp, KoboldCpp, Jan | ≈ 5 GB in Q4 |
| …-AWQ, …-GPTQ, …-FP8 repositories | Quantization designed for GPU servers | vLLM and equivalent engines | ≈ 5 to 9 GB |
| …-MLX-4bit repositories | Apple Silicon format | MLX, LM Studio on Mac | ≈ 5 GB |
Labs mostly publish the first type: the original weights exactly as they come out of training. The GGUF files desktop applications need are produced separately by converters: some labs publish their own directly, while community contributors such as bartowski, unsloth, and the ggml-org organization cover almost everything else, usually just a few hours after a highly anticipated release. To find them without searching at random, open the original model's page and follow the Quantizations link in the model tree displayed on the right side of the page, or simply search for the model name followed by GGUF in the Hub search bar, which generally lists the most-downloaded conversions first.
#Choose the right file for your card
A GGUF repository often offers a good ten or so files for the same model, one for each available quantization level, from most compressed to most faithful. The file size roughly corresponds to the VRAM the weights will occupy once the model is loaded into memory. Always keep 1 to 3 GB of VRAM free for conversation context and system operating headroom.
| Model | Q4_K_M | Q5_K_M | Q8_0 | BF16 | A card with plenty of headroom at Q4 |
|---|---|---|---|---|---|
| Qwen 3 8B | 5 GB | 6 GB | 9 GB | 16 GB | 8 GB |
| Gemma 4 12B | 7 GB | 9 GB | 13 GB | 24 GB | 12 GB |
| Qwen 3 14B | 9 GB | 11 GB | 16 GB | 28 GB | 12 GB |
| gpt-oss 20B | 13 GB | 16 GB | 23 GB | 42 GB | 16 GB |
| Mistral Small 3.2 24B | 14 GB | 17 GB | 26 GB | 48 GB | 16 GB |
| Qwen 3.8 27B | 16 GB | 19 GB | 29 GB | 54 GB | 24 GB |
| Llama 3.3 70B | 40 GB | 48 GB | 75 GB | 140 GB | 2 × 24 GB |
- Q4, Q5, Q8, or FP16: choosing your quantization
- VRAM: know yours and what it can handle
- The VRAM calculator
#Three ways to download
The simplest approach is to let your tool handle it. The LM Studio search queries Hugging Face directly and indicates which files fit on your machine. Ollama and llama.cpp can both pull a GGUF repository by name.
The --include filter matters: without it, you download every quantization in the repository, often more than 100 GB. Third option, the browser: open the Files and versions tab and click the download arrow next to the file. Convenient for a standalone GGUF, but unsuitable for split original weights.
#Licenses: open does not mean unrestricted
| License | Catalog examples | Commercial use |
|---|---|---|
| Apache 2.0 | Qwen 3, Gemma 4 12B, gpt-oss, Mistral Small 3.2 family | Yes, with a note |
| MIT | DeepSeek distillations R1, GLM 4.7 Flash | Yes |
| Llama Community License | Llama 3.3 70B | Yes, under certain conditions: naming rules, usage policy, user threshold |
| Non-commercial, research | Various research publications | No |
When in doubt, the label shown on the model page is authoritative, and a model derived, fine-tuned, or merged from another generally inherits its base model’s obligations, even when the derived model’s README says nothing explicitly. For an enterprise deployment or commercial product, these two minutes of careful license reading prevent you from building an entire integration on a model whose commercial use is actually restricted—a common and costly mistake to fix later.
#Is it safe?
- Prefer .safetensors and .gguf
- These are pure data formats. Older PyTorch .bin and .pt files rely on Python's pickle mechanism, which can execute code when loaded. The Hub flags them with a warning.
- Beware of “trust remote code”
- Some repositories, particularly those for newer architectures not yet integrated into the transformers library, include custom Python code that the loader requires you to explicitly authorize before continuing. Grant this authorization only to organizations you recognize and trust: this code runs with the same permissions as the rest of your Python program, with no special restrictions.
- See who publishes it
- The cumulative download count, the list of other repositories published by the same organization, and a link to this repository from the lab's official website are good trust signals. Impersonator accounts, using a name resembling that of a known organization with a typo or an extra underscore, exist on the Hub and elsewhere on the web: a quick look at the download count and the account's age is often enough to spot them.
- The model stays on your machine
- Once downloaded, a GGUF run by llama.cpp or Ollama makes no further network calls to Hugging Face or anyone else, unlike a model used through a cloud-hosted API. Your prompts and data never leave your machine—a key privacy argument for choosing local AI over an online service.
- Official Hugging Face Hub documentation
- The QuelLLM catalog: 249 models sorted by hardware
- Source: Wikipedia, company history
- Source: official announcement of the new hf CLI
#FAQ
What is Hugging Face used for?+
Is Hugging Face free?+
Is Hugging Face a French company?+
Do you need an account to download a model?+
Which file should you download for Ollama or LM Studio?+
Do you still need huggingface-cli to download a model?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.