Beginner 7 minConcepts

Open weights or open source? The difference that compte

“Open-source LLM” is used far too loosely. In almost every case, these models are open-weight, not open-source: you get the network weights, but neither the training code nor the data. This distinction isn't mere pedantry—it determines what you're allowed to do with the model. This guide defines the terms and provides a framework for evaluating how open a model really is before building on it.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#The initial misunderstanding

When you download Llama, Mistral, or Qwen via Ollama, you hear the term “open-source” everywhere. It is almost always a misuse of the term. The precise term is open-weights: what the lab publishes is the model’s weights—the billions of numerical parameters learned during training. Everything else (the code used for training, the data corpus, the exact recipe) usually remains closed.

The confusion comes from a tempting but false analogy: weights are equated with the model's “source code.” But weights are not readable and modifiable code like a software source codebase. They are the compiled result of training. Receiving the weights is somewhat like receiving a powerful, reconfigurable executable binary—not the complete toolchain needed to rebuild it from scratch.

i
The key takeaway
Open weights and open source are not the same thing. Nearly all the “open-source LLMs” you encounter are actually open-weight models. Understanding this difference means understanding what you can and cannot do with the model.

#Weights, code, data: the three levels of openness

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

To evaluate a model's openness, you need to distinguish three independent components. A model may open one, two, or all three. The combination determines its actual degree of openness.

The weights (weights)
The network's numerical parameters, distributed across files (often in GGUF or safetensors format). With them, you can run, quantize, and fine-tune the model. This is the most commonly open level.
The training code
The scripts, detailed architecture, hyperparameters, pretraining and alignment recipe. Without it, you can't reproduce the model—only use it as-is.
Training data
The corpus on which the model was trained. Almost never published, for legal (copyright), competitive, and practical reasons (terabytes of text).

True open-source software delivers the equivalent of all three: you can read, modify, and recompile it. An open-weights LLM generally opens only the first layer. That is useful and often sufficient for self-hosting—but it is not open source in the strict sense.

→
The reproducibility test
The deciding question: “With what has been published, can I rebuild the model from scratch?” If the answer is no because the code or data is missing, then it is not open source, regardless of the marketing terminology used.

#What open-weight really means, and what it enables

An open-weights model gives you access to the weights, generally under a license that permits local use, modification, and often commercial use. In self-hosting practice, that is exactly what you need day to day.

Run locally
Load the model in Ollama, LM Studio, or llama.cpp and run it on your machine, offline, without sending your data elsewhere.
Quantizer
Reduce the memory footprint (Q4_K_M, Q5_K_M, Q8_0…) to fit a large model on your GPU. A 14B model in Q4 fits in ~9 GB of VRAM.
Fine-tuner
Adapt the model to your domain with your own examples (LoRA, QLoRA). You start from the published weights as a base.
Redistribute
Depending on the license, you can republish a modified version—this is what fuels the thousands of community variants on Hugging Face.

In practical terms, when you launch a model with Ollama, you are using open weights. The daemon listens on http://localhost:11434 and serves a model whose training code and data you know nothing about—and for local use, that does not matter at all.

Terminal
# Télécharger et exécuter des poids open-weights en local
ollama run qwen3.5:4b

# Inspecter la licence embarquée dans le modèle
ollama show qwen3.5:4b --license
i
Check the license before building on it
The ollama show <model> --license command displays the license text provided with the weights. Make a habit of reading it before any serious project: it defines your actual rights, not the model's “open” reputation.

#Why Llama isn’t strictly open source

Meta’s Llama is the typical example of a model called “open source” even though it is not. Meta publishes the weights under a proprietary license, the Llama Community License, which is neither Apache 2.0 nor MIT. Two points disqualify it under the classic definition of open source.

A commercial threshold clause
The license imposes special conditions beyond a very large number of monthly active users. A truly open-source license does not discriminate against anyone based on size or use (the non-discrimination criterion).
Usage restrictions
The acceptable use policy prohibits certain use cases. Traditional open source does not allow restricting application domains (“no discrimination based on purpose”).
Neither code nor data
Meta publishes neither the complete training code nor the corpus. It is impossible to reproduce Llama from what is distributed.

The Open Source Initiative, which maintains the reference definition of open source, considers that this type of license does not meet the criteria. Llama is therefore, strictly speaking, an open-weights model with a permissive-but-restricted license—not an open-source model. The same reasoning applies to many models with a proprietary “community” license.

!
The trap of custom licenses
A license called “community” or “research” isn’t open-source by default. Some prohibit commercial use; others prohibit training competing models. Never assume what rights you have: read the license specific to the model you’re using.

#Truly open-source models

The good news: some models deserve the label. Two families stand out depending on their degree of openness.

#Weights under an OSI-approved license

Many models publish their weights under Apache 2.0 or MIT—recognized open-source licenses with no restrictions on use or purpose. This includes several Mistral models (such as Mistral Small 24B), the Qwen family (Qwen 3.5, Qwen 3.8) in their Apache 2.0 versions, and Gemma, which Google moved to Apache 2.0 with Gemma 4 (April 2026). You can use them commercially, modify them, and redistribute them freely. However, an Apache license on the weights does not mean that the training data or code has been published.

#Full breakdown: weights + code + data

A step above that, a few projects publish everything: weights, training code, and data. Initiatives such as OLMo (Allen Institute for AI) and the Pythia family (EleutherAI) explicitly aim for complete reproducibility. They are the only ones that are truly open-source in the fullest sense—but they are rare and don’t always match the best open-weight models in raw quality.

→
Open isn't quality
The most open model isn’t necessarily the best-performing. The best models for local use are often open-weight models (with restricted licenses) rather than fully open-source models. Choose based on your actual needs: performance, legal freedom, or reproducibility.

#A framework for evaluating openness

When faced with a new model, ask yourself these questions in order. They take you from the marketing slogan to a concrete evaluation of what you will actually be able to do with it.

  1. 01
    Are the weights public?
    Can they be downloaded without a commercial agreement or a waiting list? If so, they are at least open-weight. If not (API-only access), the model is closed, regardless of the messaging.
  2. 02
    What license is it under?
    Apache 2.0 or MIT = permissive and open-source for the weights. A proprietary “community/research” license = read the restrictions carefully (user thresholds, commercial use, training competitors).
  3. 03
    Is commercial use allowed?
    A decisive point if you're building a product. Some licenses prohibit it outright; others restrict it above a threshold.
  4. 04
    Has the training code been published?
    If it is, you can understand and reproduce the recipe. Rare outside academically oriented projects.
  5. 05
    Are the data documented?
    Published corpus or at least a described one? This is the final step toward full open source, and the one reached least often.

A model that answers “yes” to the first two questions already covers the vast majority of self-hosting needs. The last three questions matter mainly for research, auditing, and the strictest compliance requirements.

Closed
Weights unavailable, API only. Typical examples: proprietary cloud models.
Restricted open weights
Public weights, proprietary license with restrictions. Example: Llama under the Community License.
Permissive open weights
Public weights under Apache 2.0 / MIT. Examples: Mistral Small 24B, Qwen 3.5 / 3.8, Gemma 4.
Fully open source
Weights + code + data published and reproducible. Examples: OLMo, Pythia.

#What this concretely changes for you

The distinction is not theoretical. Depending on your situation, it has direct practical consequences.

You self-host for personal use
The distinction matters little to you. Open weights are more than enough: you run the model, quantize it, and your data stays with you. Choose based on performance and the size that fits your GPU.
You are building a commercial product
The license becomes critical. Check that commercial use is permitted and that no threshold will catch you out. Prefer Apache 2.0 / MIT for peace of mind.
You need to audit or prove compliance
You need the code and ideally the data. Only fully open-source models (OLMo, Pythia) provide true end-to-end traceability.
You want to fine-tune and redistribute
Make sure the license allows derivative works and their redistribution. Most permissive licenses do; some custom licenses place limits on them.
i
In summary
To run an LLM at home, open weights are almost always sufficient, and that’s what you already use. The term “open source” only becomes a real concern when the license or reproducibility determines the constraints of your project—commercial product, audit, research.

#Go further

Once the question of openness is clarified, these guides help you move on to hands-on self-hosting:

What is Ollama and how does it work
The starting point for downloading and running open-weight models locally, with the basic commands (run, pull, list).
Choose your quantization (Q4, Q5, Q8, FP16)
Once you have retrieved your weights, quantization determines their memory footprint on your card.
Run an LLM locally without a GPU (CPU only)
If you don't have a dedicated graphics card, how to run open-weight models on the CPU anyway.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.