Tokens and tokenization: understanding what a LLM
An LLM does not read words or letters: it reads tokens, fragments of text produced by a process called tokenization. This invisible unit determines everything—how much your model can memorize, how quickly it responds, and why the same text in French “weighs” more than in English. This guide explains concretely what a token is, how LLM tokenization works, and what changes when you run a model locally.
#Why talk about tokens
When you chat with a local LLM through Ollama or LM Studio, you type sentences. The model never sees your sentences as written. Before the first computation, your text is transformed into a sequence of numbers, each representing a token. Everything the model does—understanding and generating—happens at the token level, not the word level.
Understanding tokens is not an academic detail. It explains three very concrete things: why a document does or does not fit in the context window, why your model generates at a given speed (the famous tokens per second), and why cloud API billing or a context limit is always measured in tokens, never words.
#What exactly is a token?
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A token is a fragment of text: sometimes a whole word, often part of a word, and sometimes a single character or punctuation mark. It is neither a letter nor a word in the strict sense—it is a statistical unit chosen by the tokenization algorithm to represent text as efficiently as possible.
The most useful rule of thumb: in English, 1 token ≈ 4 characters ≈ 0.75 words. In other words, 100 tokens amount to about 75 English words. For French, the ratio is significantly less favorable; we’ll return to this below.
- "chat"
- A common, short word: often just 1 token.
- "anticonstitutionnellement"
- A rare, long word: split into several tokens (anti / constitution / nelle / ment…).
- " " (space)
- The space is generally attached to the beginning of the next token, not isolated.
- "123456"
- Numbers are often split digit by digit or into small groups.
- 😀
- A single emoji can cost several tokens.
Very common words inherit a single token because they appear everywhere in the training data. Rare, technical, or underrepresented-language words are reconstructed from subpieces, which makes them more "expensive" in tokens.
#How text is actually split
Most modern LLMs use an algorithm family called BPE (Byte Pair Encoding) or one of its variants (WordPiece, Unigram). The principle is to start with raw characters, then progressively merge the most frequent symbol pairs in the training corpus until reaching a fixed-size vocabulary—typically 32,000 to 200,000 tokens, depending on the model.
- 01Initial breakdownThe text is reduced to its basic bytes or characters. Nothing is lost: any string can be represented.
- 02Learned mergesThe tokenizer applies the list of merges learned during training (e.g., “t” + “ion” → “tion”), in frequency order.
- 03Conversion to identifiersEach final token is replaced by its number in the vocabulary. The model manipulates only these integers.
Important consequence: the vocabulary is fixed during training. A model trained primarily on English will have a vocabulary optimized for English and will split French into smaller, more numerous pieces. This is the root of the French-language overhead.
#Why French costs more than English
With equivalent content, a French text commonly uses 15 to 30% more tokens than its English translation. On some heavily English-focused models, the gap can exceed 50%. Three reasons add up.
- Unbalanced vocabulary
- Tokenizers are trained mostly on English: common English words have their own token, but French words do not.
- Accents and special characters
- é, è, à, ç, œ… are rarer in the vocabulary and are sometimes split into multiple tokens (or even bytes).
- Richer morphology
- Conjugations, agreements, and elisions (l’, d’, qu’) multiply the forms of the same word, which are less well covered by the vocabulary.
Concrete example: the English sentence “The cat is on the table” is about 6 tokens. Its French version, “Le chat est sur la table,” is closer to 7 to 8 depending on the model. Over an entire paragraph, the difference becomes significant—and you pay for it twice: in context-window space and generation time.
#Tokens and context window
A model's context window is measured in tokens, not words or characters. A model advertised with a 32,768-token context can “see” the equivalent of ~24,000 English words at any given time—but only ~18,000 to 20,000 French words because of the tokenization overhead.
This window includes everything: the system prompt, conversation history, your current message, pasted documents, and the response currently being generated. When the total exceeds the limit, the model truncates—usually the oldest content—and "forgets" the beginning of the exchange.
- System prompt
- Counted on every call. A verbose system prompt constantly eats into the context.
- History
- Each conversation turn accumulates. A long discussion eventually fills the window.
- Documents (RAG, copy-paste)
- A 10-page PDF can quickly amount to several thousand tokens.
- Generated response
- The output also takes up space: you need to reserve enough to generate a response.
#Local tokens and speed
LLM speed is measured in tokens per second (tok/s). This is the universal benchmark unit. Two moments must be distinguished, and they are often confused.
- Prompt / prefill
- The time it takes to “read” your entire prompt. The more input tokens there are, the longer it takes for the first response to start.
- Generation / decode
- The rate at which response tokens appear, one by one. This is the tok/s you see on screen.
The direct consequence for French: because the same content represents more tokens, your model mechanically takes longer to process a French-language prompt and generate an equivalent-length French-language response. The impression that “it's slower in French” isn't subjective—it is token counting.
Decode speed depends mainly on the model and hardware (active parameters per token, memory bandwidth, quantization). But with fixed hardware, reducing the number of input tokens — a shorter system prompt, pruned history — noticeably speeds up time-to-first-token.
#Count a text's tokens yourself
The best way to build intuition is to measure. Here are three approaches, from simplest to most precise.
#Estimate roughly
For a quick estimate without installing anything: divide the character count by 4 for English, or by 3 to 3.5 for French. It's rough but sufficient to determine whether a document fits in a context window.
#Count in Python with the actual tokenizer
For an exact count, use the model's tokenizer. Hugging Face's tokenizers / transformers library loads the actual tokenizer for an open-weight model:
Running this script on your own text is the most revealing exercise: you can see for yourself which French words get split apart and how much more compact English is.
#Read the count from the Ollama API
Ollama already exposes token counts in its responses. The daemon listens by default on http://localhost:11434; an API call returns prompt_eval_count (prompt tokens) and eval_count (generated tokens):
You’ll also find eval_duration: divide eval_count by the duration to get your actual throughput in tokens per second, on your machine, with your quantization.
#Save tokens without sacrificing quality
Since every token costs context and speed, a few simple habits make a real difference, especially in French.
- Concise system prompt
- It is sent with every call. Every unnecessary sentence costs you on every turn.
- Trim history
- Summarize or cut off old exchanges instead of dragging the entire conversation along.
- RAG instead of pasting everything
- Inject only the relevant passages from a document, not the entire document.
- Targeting output length
- Requesting a short response generates fewer tokens—and therefore responds faster.
#Go further
Tokens are the thread connecting several basic concepts. Three guides build directly on this one: the first details the context window that tokens fill, the second explains how quantization affects speed measured in tokens per second, and the third shows how the transformer processes those tokens internally.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.