Beginner 9 minBasics

LLM: what is it? Simple definition and fonctionnement

What exactly is an LLM? In one sentence: a program that has read a gigantic amount of text and, based on what it learned, predicts the most likely next word—again and again, until it forms a response. This guide explains without jargon how a large language model is trained, how it generates text token by token, and the difference between a cloud LLM (ChatGPT, Gemini) and a local LLM that you can run on your own machine.

By Mohamed Meguedmi·Update 2026-09-02·Tested on Windows, macOS, and Linux

#What is an LLM? The simple definition

LLM is an acronym for Large Language Model, in French “grand modèle de langage.” The word “large” refers to size: these models contain billions of parameters, a kind of numerical settings adjusted during training. “Language model” means that the program models language: it learned how likely a given sequence of words is in a given situation.

In concrete terms, an LLM does just one thing: given a piece of text, it predicts what comes next. You write “The capital of France is,” and it calculates that the most likely word after that sentence is “Paris.” All the apparent magic of an assistant like ChatGPT follows from repeating this mechanism thousands of times: drafting an email, summarizing a document, or writing code is nothing more than a long sequence of “what is the most likely next word?”

i
In one image
Imagine your phone's predictive typing, but trained on thousands of billions of words instead of your text messages. An LLM is that predictive keyboard taken to the extreme—powerful enough to hold a coherent conversation.

#Tokens: how an LLM reads and writes

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

An LLM doesn’t see words or letters the way we do. It works with tokens: fragments of text. A token can be a complete short word (“cat”), part of a word (“anti”, “constitution”), a punctuation mark, or a space. On average, in French, a token corresponds to roughly 3 or 4 characters, or about ¾ of a word.

Before processing your message, the model breaks it into tokens (this is called tokenization), then converts each token into a list of numbers. Everything it computes, it computes using those numbers. When it responds, it generates one token at a time, then converts them back into readable text. That is why we talk about generation “token by token.”

Token
The basic unit handled by the model: a short word, part of a word, or a symbol.
Context
The number of tokens the model can “keep in mind” at once (question + history + answer). This is called the context window, often 4,000, 32,000, or even hundreds of thousands of tokens.
Practical reference point
A page of text ≈ 500 to 700 tokens. 1,000 tokens ≈ 750 words. Cloud billing and local speed are both measured in tokens.
→
Why it matters
The context window is a physical limit: if your document exceeds the context size, the model sees only part of it. Understanding tokens helps explain why an LLM "forgets" the beginning of a very long conversation.

#Training: how an LLM learns

An LLM is not programmed rule by rule. It is trained. It is shown immense volumes of text (web pages, books, code, articles) and asked, at each point, to guess the next word. At first, it is almost always wrong. With each error, an algorithm slightly adjusts its billions of parameters so the next prediction is a little better. Repeated billions of times, this process produces a statistical understanding of grammar, facts, and reasoning.

This phase takes place in two main steps. Pre-training builds the model's general knowledge from raw text. Then fine-tuning and alignment teach it to follow instructions and respond helpfully and politely, often with the help of human feedback. An “instruct” or “chat” model has gone through this second step; that's the one used for conversation.

Parameters (the “billions”)
The internal settings adjusted during training. A “7B” model has 7 billion parameters; a “70B” has 70 billion. The more it has, the more capable it is—and the more memory it requires.
Weights
Another name for the parameters once they are frozen. Downloading a model means downloading its weights: a large file several gigabytes in size.
Training data
The text read during training. The model does not “store” it word for word; it retains statistical patterns.
i
Training ≠ usage
Training costs millions of euros and weeks of computation on thousands of GPUs; no one does it at home. Once the weights are published, however, using them (inference) runs on an ordinary PC. That is the whole promise of the local LLM.

#Inference: generating text one token at a time

Once trained, using the model is called inference. You send it text (the prompt), and the model produces its response one token at a time. Here is the exact loop: it reads all available text, calculates the most likely next token, appends it, rereads everything—the prompt plus the new token—and starts over. It stops when it generates a special end token or reaches the configured limit.

That is why, in a chat interface, the response appears word by word in real time: you are literally watching the generation happen. Speed is measured in tokens per second. With a local LLM, it depends on your hardware — a good GPU generates dozens of tokens per second, while a CPU alone generates far fewer.

The model does not always rigidly choose the most probable token. A setting called temperature introduces a degree of controlled randomness: at low temperature, responses are predictable and factual; at high temperature, they are more creative and varied. This same mechanism explains why an LLM can give two different answers to the same question.

!
The hallucination trap
An LLM predicts what is plausible, not what is true. When it doesn’t “know,” it still generates the most probable continuation—which may be false but stated confidently. This is called a hallucination: always keep it in mind for facts, figures, and citations.

#What an LLM is not

Understanding the mechanics prevents common misunderstandings. An LLM isn't a database: it doesn't retrieve a stored answer; it reconstructs one statistically. It isn't connected to the internet by default: it only knows what it read up to its training date, unless you connect a search tool or RAG to it. And it doesn't “understand” in the human sense: it models language extremely well, which is enough to make it useful, but it remains a prediction, not consciousness.


#Cloud LLM vs. local LLM: the real difference

There are two ways to use an LLM. With a cloud LLM, the model runs on a company's servers (OpenAI for ChatGPT, Google for Gemini, Anthropic for Claude). You send your text over the internet, their GPUs compute the response, and send it back. You download nothing; in exchange, your data passes through a third party, and you depend on a subscription and an internet connection.

With a local LLM, you download the weights of an open-weight model (Llama, Qwen, Mistral, Gemma…) and run it on your own computer. Inference takes place entirely on your machine: no data leaves, there’s no subscription, and it even works offline. The tradeoff is that you need sufficient hardware, and the best open models often remain a step behind the largest proprietary cloud models.

Privacy
Cloud: your prompts are sent to the provider. Local: everything stays on your machine.
Cost
Cloud: monthly subscription or pay per token. Local: free to use once you've acquired the hardware.
Offline
Cloud: internet connection required. Local: works without internet once the model has been downloaded.
Power
Cloud: access to the largest models without hardware. Local: limited by your GPU/RAM, but highly capable with even a 12 GB card.
Control
Cloud: the provider can change the model or the rules. Local: the model belongs to you and does not change without your approval.

#Why local LLMs exist

If the cloud is so convenient, why run an LLM at home? Three main reasons. Privacy comes first: medical documents, contracts, proprietary code, personal notes—many people and companies refuse to send this content to third-party servers. Independence comes next: no subscription, no message limits, and no disruption if the service changes its pricing or disappears. And technical control: you can tune a local model, integrate it with your own tools, and use it on a plane or in an area without network access.

What made all this possible is quantization: a technique that compresses model weights so they fit in memory without sacrificing too much quality. Thanks to it, a model that once required a server can now run on a gaming PC. The Q4_K_M format is the recommended compromise for getting started; Q5_K_M, Q8_0, and FP16 offer more precision at the cost of more memory.

i
VRAM benchmark for Q4
How much memory for which model? 3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. A RTX 3060 12 GB comfortably runs a 7B or 14B; a RTX 4090 24 GB or a Mac Apple Silicon 48 GB targets larger models.

#Try an LLM at home in 10 minutes

The best way to understand an LLM is to run one. The simplest stack consists of two pieces: Ollama, a daemon that downloads and runs models (it listens on http://localhost:11434), and an interface such as Open WebUI or LM Studio for chatting in a chat window. Here is the minimal command-line workflow, using a small model that runs even without a dedicated GPU.

  1. 01
    Install Ollama
    Download the installer from ollama.com for Windows, macOS, or Linux, and run it. Once installed, Ollama runs in the background and exposes its local API on port 11434.
  2. 02
    Download and run a model
    A single command downloads the weights and then opens the chat directly in the terminal. The first launch downloads a few gigabytes; subsequent launches are instant.
  3. 03
    Discuss
    Ask a question and watch the response appear token by token: you can see inference in real time. Type /bye to leave the conversation.
  4. 04
    Verify that everything is local
    Turn off Wi-Fi and ask another question. The model keeps responding: proof that the computation is happening 100% on your machine.
Terminal
# Lancer un petit modèle rapide (≈2 Go, tourne même sans GPU)
ollama run llama3.2:3b

# Pour un modèle plus capable si vous avez ~5 Go de VRAM
ollama run qwen3

# Vérifier les modèles installés et que le daemon répond
ollama list
curl http://localhost:11434/api/tags
→
Tag names evolve
The exact tags (llama3.2:3b, qwen3…) and available sizes change regularly in the Ollama library. Check ollama.com/library for the precise name and size in GB of each variant before downloading.

#Go further

You now know what an LLM is and how it works. These guides build on each concept covered here:

Get started with Ollama
“What is Ollama and how does it work?” covers installation, basic commands (run, pull, list), and the hardware required to get started.
Understanding tokens
“Understanding the context window” explains how the model “sees” your messages and why it forgets long conversations.
Choose your quantization
“Choosing your quantization (Q4, Q5, Q8, FP16)” helps you find the right balance between quality and memory for your card.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.