Beginner 11 minConcepts

Hugging Face: what it is and how to use it ?

Direct response

Hugging Face is the platform where almost all open AI models are published: Qwen, Llama, Gemma, Mistral, gpt-oss, DeepSeek. It is often described as the GitHub of machine learning: the repository where the files come from, including those that Ollama and LM Studio download for you. The Hub passed three million public models on August 18, 2026, without any editorial review.

Hugging Face is the platform where almost all open AI models are published: Qwen, Llama, Gemma, Mistral, gpt-oss, DeepSeek. It’s often described as the GitHub of machine learning. If you run AI on your own machine, it’s the repository all the files come from, including the files that Ollama and LM Studio download for you. The Hub officially passed the three-million-public-model mark on August 18, 2026, at a rate of about 2,700 to 3,000 new models per day, with no filtering at all. This page teaches you how to read it: which repository to open, which file to choose for your graphics card, and what to check before clicking.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What is Hugging Face?

Hugging Face is a company founded in 2016 by three French entrepreneurs, Clément Delangue, Julien Chaumond, and Thomas Wolf, and based between New York and Paris, with its headquarters remaining in the United States. It started with a chatbot app for teenagers, a project now abandoned, hence the name borrowed from the emoji that smiles with its hands open. It then turned to open-source tooling for machine learning: its transformers library became the de facto standard way to load and run a language model in Python, adopted by most research labs and companies in the industry. The company eventually built the Hub, a hosting platform based on the Git version-control system where anyone can publish models, datasets, and small web demos for free, without prior editorial review.

For local AI, the Hub acts like an app store, with two key differences: almost everything there is free, and nothing is selected for you. Unlike a traditional app store, nobody tests, approves, or ranks models before publication: anyone can create a free account and upload a repository in a few minutes, without any validation or quality control whatsoever. This complete lack of filtering is both the Hub's strength—everything arrives there first, often even before a lab's official announcement—and its main pitfall for beginners, who may face hundreds of search results without knowing which one is reliable. Knowing how to navigate it, distinguish an official repository from a third-party copy and a serious conversion from an abandoned experiment, is the skill to acquire before downloading anything.

#The Hub’s three sections

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
RadiusContentInterest in local AI
ModelsWeight files, configuration, documentationThis is where the file you'll run is located
DatasetsTraining and evaluation dataOnly if you’re doing fine-tuning or testing
SpacesSmall web applications, mostly demosTry a model in the browser before downloading 15 GB

#Read a model page

Each model lives at an address in the form huggingface.co/organisation/nom-du-modele, for example huggingface.co/Qwen/Qwen3-8B. Four distinct elements on this page deserve your attention before any download.

The organization name
Qwen, google, meta-llama, mistralai, openai, deepseek-ai are the labs themselves, verified. Any other name refers to a third party: sometimes excellent, but never the lab's original.
The model card
The README file: what the model is, how it was trained, its intended uses, its scores, its prompt format, and its limitations. Quality ranges from highly comprehensive to empty.
The license
Displayed prominently at the top of the model page, next to the model name. See the dedicated section below for what each common license actually allows.
The Files and versions tab
The complete list of downloadable files. A quick glance is enough to immediately tell what type of repository you are looking at.

#Which repository to choose: original weights or GGUF

What you see in FilesWhat it isFor which toolsSize for an 8B model
model-00001-of-00004.safetensors…Original weights, in BF16, split across multiple filestransformers, vLLM, fine-tuning tools≈ 16 GB
…-Q4_K_M.gguf, …-Q8_0.ggufQuantized conversions, one file per levelOllama, LM Studio, llama.cpp, KoboldCpp, Jan≈ 5 GB in Q4
…-AWQ, …-GPTQ, …-FP8 repositoriesQuantization designed for GPU serversvLLM and equivalent engines≈ 5 to 9 GB
…-MLX-4bit repositoriesApple Silicon formatMLX, LM Studio on Mac≈ 5 GB

Labs mostly publish the first type: the original weights exactly as they come out of training. The GGUF files desktop applications need are produced separately by converters: some labs publish their own directly, while community contributors such as bartowski, unsloth, and the ggml-org organization cover almost everything else, usually just a few hours after a highly anticipated release. To find them without searching at random, open the original model's page and follow the Quantizations link in the model tree displayed on the right side of the page, or simply search for the model name followed by GGUF in the Hub search bar, which generally lists the most-downloaded conversions first.

#Choose the right file for your card

A GGUF repository often offers a good ten or so files for the same model, one for each available quantization level, from most compressed to most faithful. The file size roughly corresponds to the VRAM the weights will occupy once the model is loaded into memory. Always keep 1 to 3 GB of VRAM free for conversation context and system operating headroom.

Calculated from the QuelLLM catalog · 20/09/2026 · sizes rounded up to the nearest GB
ModelQ4_K_MQ5_K_MQ8_0BF16A card with plenty of headroom at Q4
Qwen 3 8B5 GB6 GB9 GB16 GB8 GB
Gemma 4 12B7 GB9 GB13 GB24 GB12 GB
Qwen 3 14B9 GB11 GB16 GB28 GB12 GB
gpt-oss 20B13 GB16 GB23 GB42 GB16 GB
Mistral Small 3.2 24B14 GB17 GB26 GB48 GB16 GB
Qwen 3.8 27B16 GB19 GB29 GB54 GB24 GB
Llama 3.3 70B40 GB48 GB75 GB140 GB2 × 24 GB

#Three ways to download

The simplest approach is to let your tool handle it. The LM Studio search queries Hugging Face directly and indicates which files fit on your machine. Ollama and llama.cpp can both pull a GGUF repository by name.

From Ollama or llama.cpp
# Ollama : lancer n'importe quel dépôt GGUF du Hub
ollama run hf.co/bartowski/Qwen_Qwen3-8B-GGUF:Q4_K_M

# llama.cpp : télécharger et servir
llama-server -hf bartowski/Qwen_Qwen3-8B-GGUF:Q4_K_M
With the command-line tool
pip install -U huggingface_hub
hf download bartowski/Qwen_Qwen3-8B-GGUF --include "*Q4_K_M.gguf" --local-dir ./models

The --include filter matters: without it, you download every quantization in the repository, often more than 100 GB. Third option, the browser: open the Files and versions tab and click the download arrow next to the file. Convenient for a standalone GGUF, but unsuitable for split original weights.

i
Does an old tutorial tell you to use huggingface-cli?
Ignore it: the huggingface-cli download command is deprecated and has been completely removed since version 1.0 of the huggingface_hub library, replaced by the shorter hf command, which Hugging Face itself describes as “a faster, more user-friendly CLI.” The replacement is mechanical: use hf download wherever the old tutorial used huggingface-cli download, with the same options.
i
Models with conditional access
Some labs, including Meta for the Llama family, require you to accept their terms before downloading. Create a free account, click the access button on the model page, then authenticate the command-line tool once with hf auth login and a token created in your account settings. Community GGUF conversions are generally not locked, but the original license still applies to you.

#Licenses: open does not mean unrestricted

Licenses recorded in the QuelLLM catalog · 20/09/2026
LicenseCatalog examplesCommercial use
Apache 2.0Qwen 3, Gemma 4 12B, gpt-oss, Mistral Small 3.2 familyYes, with a note
MITDeepSeek distillations R1, GLM 4.7 FlashYes
Llama Community LicenseLlama 3.3 70BYes, under certain conditions: naming rules, usage policy, user threshold
Non-commercial, researchVarious research publicationsNo

When in doubt, the label shown on the model page is authoritative, and a model derived, fine-tuned, or merged from another generally inherits its base model’s obligations, even when the derived model’s README says nothing explicitly. For an enterprise deployment or commercial product, these two minutes of careful license reading prevent you from building an entire integration on a model whose commercial use is actually restricted—a common and costly mistake to fix later.

#Is it safe?

Prefer .safetensors and .gguf
These are pure data formats. Older PyTorch .bin and .pt files rely on Python's pickle mechanism, which can execute code when loaded. The Hub flags them with a warning.
Beware of “trust remote code”
Some repositories, particularly those for newer architectures not yet integrated into the transformers library, include custom Python code that the loader requires you to explicitly authorize before continuing. Grant this authorization only to organizations you recognize and trust: this code runs with the same permissions as the rest of your Python program, with no special restrictions.
See who publishes it
The cumulative download count, the list of other repositories published by the same organization, and a link to this repository from the lab's official website are good trust signals. Impersonator accounts, using a name resembling that of a known organization with a typo or an extra underscore, exist on the Hub and elsewhere on the web: a quick look at the download count and the account's age is often enough to spot them.
The model stays on your machine
Once downloaded, a GGUF run by llama.cpp or Ollama makes no further network calls to Hugging Face or anyone else, unlike a model used through a cloud-hosted API. Your prompts and data never leave your machine—a key privacy argument for choosing local AI over an online service.

#FAQ

What is Hugging Face used for?+
For publishing and downloading open AI models, datasets, and small web demos. For local AI, it is the main source of model files, including quantized GGUF versions loaded directly by Ollama, LM Studio, and llama.cpp, often produced by community contributors a few hours after the original lab officially releases a model.
Is Hugging Face free?+
Browsing and downloading public models is completely free and requires no account, except for gated-access models, for which an account is still free. The company actually charges for hosted inference on its servers, additional computing power for Spaces, private repository storage, and its enterprise offerings.
Is Hugging Face a French company?+
It was founded in 2016 by three French entrepreneurs, Clément Delangue, Julien Chaumond, and Thomas Wolf, and still maintains a substantial team in Paris, but its headquarters are in New York, United States. It is one of the few truly major global AI platforms with such strong French roots.
Do you need an account to download a model?+
Not in the vast majority of cases: browsing and downloading public models requires nothing. A free account and an access token are only needed for models whose lab requires prior acceptance of its terms of use, such as Meta’s Llama family or some Google releases.
Which file should you download for Ollama or LM Studio?+
A single .gguf file from a GGUF repository, generally the Q4_K_M variant to start with, whose size remains 1 to 3 GB below your graphics card’s memory capacity. The split .safetensors files from the lab’s original repository are intended for Python frameworks such as transformers and server engines such as vLLM, not desktop applications.
Do you still need huggingface-cli to download a model?+
No: this command is deprecated and was removed in huggingface_hub version 1.0. The current command, described by Hugging Face as faster and simpler, is called hf: hf download followed by the repository name works exactly where an old tutorial used huggingface-cli download, with the same filtering options.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.