LLM: what is it? Simple definition and fonctionnement
What exactly is an LLM? In one sentence: a program that has read a gigantic amount of text and, based on what it learned, predicts the most likely next word—again and again, until it forms a response. This guide explains without jargon how a large language model is trained, how it generates text token by token, and the difference between a cloud LLM (ChatGPT, Gemini) and a local LLM that you can run on your own machine.
#What is an LLM? The simple definition
LLM is an acronym for Large Language Model, in French “grand modèle de langage.” The word “large” refers to size: these models contain billions of parameters, a kind of numerical settings adjusted during training. “Language model” means that the program models language: it learned how likely a given sequence of words is in a given situation.
In concrete terms, an LLM does just one thing: given a piece of text, it predicts what comes next. You write “The capital of France is,” and it calculates that the most likely word after that sentence is “Paris.” All the apparent magic of an assistant like ChatGPT follows from repeating this mechanism thousands of times: drafting an email, summarizing a document, or writing code is nothing more than a long sequence of “what is the most likely next word?”
#Tokens: how an LLM reads and writes
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
An LLM doesn’t see words or letters the way we do. It works with tokens: fragments of text. A token can be a complete short word (“cat”), part of a word (“anti”, “constitution”), a punctuation mark, or a space. On average, in French, a token corresponds to roughly 3 or 4 characters, or about ¾ of a word.
Before processing your message, the model breaks it into tokens (this is called tokenization), then converts each token into a list of numbers. Everything it computes, it computes using those numbers. When it responds, it generates one token at a time, then converts them back into readable text. That is why we talk about generation “token by token.”
- Token
- The basic unit handled by the model: a short word, part of a word, or a symbol.
- Context
- The number of tokens the model can “keep in mind” at once (question + history + answer). This is called the context window, often 4,000, 32,000, or even hundreds of thousands of tokens.
- Practical reference point
- A page of text ≈ 500 to 700 tokens. 1,000 tokens ≈ 750 words. Cloud billing and local speed are both measured in tokens.
#Training: how an LLM learns
An LLM is not programmed rule by rule. It is trained. It is shown immense volumes of text (web pages, books, code, articles) and asked, at each point, to guess the next word. At first, it is almost always wrong. With each error, an algorithm slightly adjusts its billions of parameters so the next prediction is a little better. Repeated billions of times, this process produces a statistical understanding of grammar, facts, and reasoning.
This phase takes place in two main steps. Pre-training builds the model's general knowledge from raw text. Then fine-tuning and alignment teach it to follow instructions and respond helpfully and politely, often with the help of human feedback. An “instruct” or “chat” model has gone through this second step; that's the one used for conversation.
- Parameters (the “billions”)
- The internal settings adjusted during training. A “7B” model has 7 billion parameters; a “70B” has 70 billion. The more it has, the more capable it is—and the more memory it requires.
- Weights
- Another name for the parameters once they are frozen. Downloading a model means downloading its weights: a large file several gigabytes in size.
- Training data
- The text read during training. The model does not “store” it word for word; it retains statistical patterns.
#Inference: generating text one token at a time
Once trained, using the model is called inference. You send it text (the prompt), and the model produces its response one token at a time. Here is the exact loop: it reads all available text, calculates the most likely next token, appends it, rereads everything—the prompt plus the new token—and starts over. It stops when it generates a special end token or reaches the configured limit.
That is why, in a chat interface, the response appears word by word in real time: you are literally watching the generation happen. Speed is measured in tokens per second. With a local LLM, it depends on your hardware — a good GPU generates dozens of tokens per second, while a CPU alone generates far fewer.
The model does not always rigidly choose the most probable token. A setting called temperature introduces a degree of controlled randomness: at low temperature, responses are predictable and factual; at high temperature, they are more creative and varied. This same mechanism explains why an LLM can give two different answers to the same question.
#What an LLM is not
Understanding the mechanics prevents common misunderstandings. An LLM isn't a database: it doesn't retrieve a stored answer; it reconstructs one statistically. It isn't connected to the internet by default: it only knows what it read up to its training date, unless you connect a search tool or RAG to it. And it doesn't “understand” in the human sense: it models language extremely well, which is enough to make it useful, but it remains a prediction, not consciousness.
#Cloud LLM vs. local LLM: the real difference
There are two ways to use an LLM. With a cloud LLM, the model runs on a company's servers (OpenAI for ChatGPT, Google for Gemini, Anthropic for Claude). You send your text over the internet, their GPUs compute the response, and send it back. You download nothing; in exchange, your data passes through a third party, and you depend on a subscription and an internet connection.
With a local LLM, you download the weights of an open-weight model (Llama, Qwen, Mistral, Gemma…) and run it on your own computer. Inference takes place entirely on your machine: no data leaves, there’s no subscription, and it even works offline. The tradeoff is that you need sufficient hardware, and the best open models often remain a step behind the largest proprietary cloud models.
- Privacy
- Cloud: your prompts are sent to the provider. Local: everything stays on your machine.
- Cost
- Cloud: monthly subscription or pay per token. Local: free to use once you've acquired the hardware.
- Offline
- Cloud: internet connection required. Local: works without internet once the model has been downloaded.
- Power
- Cloud: access to the largest models without hardware. Local: limited by your GPU/RAM, but highly capable with even a 12 GB card.
- Control
- Cloud: the provider can change the model or the rules. Local: the model belongs to you and does not change without your approval.
#Why local LLMs exist
If the cloud is so convenient, why run an LLM at home? Three main reasons. Privacy comes first: medical documents, contracts, proprietary code, personal notes—many people and companies refuse to send this content to third-party servers. Independence comes next: no subscription, no message limits, and no disruption if the service changes its pricing or disappears. And technical control: you can tune a local model, integrate it with your own tools, and use it on a plane or in an area without network access.
What made all this possible is quantization: a technique that compresses model weights so they fit in memory without sacrificing too much quality. Thanks to it, a model that once required a server can now run on a gaming PC. The Q4_K_M format is the recommended compromise for getting started; Q5_K_M, Q8_0, and FP16 offer more precision at the cost of more memory.
#Try an LLM at home in 10 minutes
The best way to understand an LLM is to run one. The simplest stack consists of two pieces: Ollama, a daemon that downloads and runs models (it listens on http://localhost:11434), and an interface such as Open WebUI or LM Studio for chatting in a chat window. Here is the minimal command-line workflow, using a small model that runs even without a dedicated GPU.
- 01Install OllamaDownload the installer from ollama.com for Windows, macOS, or Linux, and run it. Once installed, Ollama runs in the background and exposes its local API on port 11434.
- 02Download and run a modelA single command downloads the weights and then opens the chat directly in the terminal. The first launch downloads a few gigabytes; subsequent launches are instant.
- 03DiscussAsk a question and watch the response appear token by token: you can see inference in real time. Type /bye to leave the conversation.
- 04Verify that everything is localTurn off Wi-Fi and ask another question. The model keeps responding: proof that the computation is happening 100% on your machine.
#Go further
You now know what an LLM is and how it works. These guides build on each concept covered here:
- Get started with Ollama
- “What is Ollama and how does it work?” covers installation, basic commands (run, pull, list), and the hardware required to get started.
- Understanding tokens
- “Understanding the context window” explains how the model “sees” your messages and why it forgets long conversations.
- Choose your quantization
- “Choosing your quantization (Q4, Q5, Q8, FP16)” helps you find the right balance between quality and memory for your card.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.