Beginner 9 minBasics

Tokens and tokenization: understanding what a LLM

An LLM does not read words or letters: it reads tokens, fragments of text produced by a process called tokenization. This invisible unit determines everything—how much your model can memorize, how quickly it responds, and why the same text in French “weighs” more than in English. This guide explains concretely what a token is, how LLM tokenization works, and what changes when you run a model locally.

By Mohamed Meguedmi·Update 2026-09-13·Tested on Windows, macOS, and Linux

#Why talk about tokens

When you chat with a local LLM through Ollama or LM Studio, you type sentences. The model never sees your sentences as written. Before the first computation, your text is transformed into a sequence of numbers, each representing a token. Everything the model does—understanding and generating—happens at the token level, not the word level.

Understanding tokens is not an academic detail. It explains three very concrete things: why a document does or does not fit in the context window, why your model generates at a given speed (the famous tokens per second), and why cloud API billing or a context limit is always measured in tokens, never words.

i
The key takeaway
LLM tokenization is the step that breaks your text into tokens before processing. It is deterministic preprocessing: the same text always produces the same tokens for a given model.

#What exactly is a token?

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A token is a fragment of text: sometimes a whole word, often part of a word, and sometimes a single character or punctuation mark. It is neither a letter nor a word in the strict sense—it is a statistical unit chosen by the tokenization algorithm to represent text as efficiently as possible.

The most useful rule of thumb: in English, 1 token ≈ 4 characters ≈ 0.75 words. In other words, 100 tokens amount to about 75 English words. For French, the ratio is significantly less favorable; we’ll return to this below.

"chat"
A common, short word: often just 1 token.
"anticonstitutionnellement"
A rare, long word: split into several tokens (anti / constitution / nelle / ment…).
" " (space)
The space is generally attached to the beginning of the next token, not isolated.
"123456"
Numbers are often split digit by digit or into small groups.
😀
A single emoji can cost several tokens.

Very common words inherit a single token because they appear everywhere in the training data. Rare, technical, or underrepresented-language words are reconstructed from subpieces, which makes them more "expensive" in tokens.

#How text is actually split

Most modern LLMs use an algorithm family called BPE (Byte Pair Encoding) or one of its variants (WordPiece, Unigram). The principle is to start with raw characters, then progressively merge the most frequent symbol pairs in the training corpus until reaching a fixed-size vocabulary—typically 32,000 to 200,000 tokens, depending on the model.

  1. 01
    Initial breakdown
    The text is reduced to its basic bytes or characters. Nothing is lost: any string can be represented.
  2. 02
    Learned merges
    The tokenizer applies the list of merges learned during training (e.g., “t” + “ion” → “tion”), in frequency order.
  3. 03
    Conversion to identifiers
    Each final token is replaced by its number in the vocabulary. The model manipulates only these integers.

Important consequence: the vocabulary is fixed during training. A model trained primarily on English will have a vocabulary optimized for English and will split French into smaller, more numerous pieces. This is the root of the French-language overhead.

→
Every model has its own tokenizer
The same text doesn't produce the same number of tokens on Llama, Qwen, Mistral, or Gemma. A count is always relative to a specific model. For a reliable count, use the tokenizer for the model you're actually running.

#Why French costs more than English

With equivalent content, a French text commonly uses 15 to 30% more tokens than its English translation. On some heavily English-focused models, the gap can exceed 50%. Three reasons add up.

Unbalanced vocabulary
Tokenizers are trained mostly on English: common English words have their own token, but French words do not.
Accents and special characters
é, è, à, ç, œ… are rarer in the vocabulary and are sometimes split into multiple tokens (or even bytes).
Richer morphology
Conjugations, agreements, and elisions (l’, d’, qu’) multiply the forms of the same word, which are less well covered by the vocabulary.

Concrete example: the English sentence “The cat is on the table” is about 6 tokens. Its French version, “Le chat est sur la table,” is closer to 7 to 8 depending on the model. Over an entire paragraph, the difference becomes significant—and you pay for it twice: in context-window space and generation time.

i
Good news: it's improving
Recent multilingual models (Qwen, Gemma, Mistral) have much better-balanced tokenizers than the first generations. The French/English gap remains, but it has narrowed significantly in models designed as multilingual from the outset.

#Tokens and context window

A model's context window is measured in tokens, not words or characters. A model advertised with a 32,768-token context can “see” the equivalent of ~24,000 English words at any given time—but only ~18,000 to 20,000 French words because of the tokenization overhead.

This window includes everything: the system prompt, conversation history, your current message, pasted documents, and the response currently being generated. When the total exceeds the limit, the model truncates—usually the oldest content—and "forgets" the beginning of the exchange.

System prompt
Counted on every call. A verbose system prompt constantly eats into the context.
History
Each conversation turn accumulates. A long discussion eventually fills the window.
Documents (RAG, copy-paste)
A 10-page PDF can quickly amount to several thousand tokens.
Generated response
The output also takes up space: you need to reserve enough to generate a response.
!
The context trap that inflates VRAM usage
Locally, expanding the context window is not free: the KV cache grows with the token count and consumes additional VRAM beyond the model weights. A 32k context can require several extra gigabytes. Too much context can overflow your GPU and send performance crashing.

#Local tokens and speed

LLM speed is measured in tokens per second (tok/s). This is the universal benchmark unit. Two moments must be distinguished, and they are often confused.

Prompt / prefill
The time it takes to “read” your entire prompt. The more input tokens there are, the longer it takes for the first response to start.
Generation / decode
The rate at which response tokens appear, one by one. This is the tok/s you see on screen.

The direct consequence for French: because the same content represents more tokens, your model mechanically takes longer to process a French-language prompt and generate an equivalent-length French-language response. The impression that “it's slower in French” isn't subjective—it is token counting.

Decode speed depends mainly on the model and hardware (active parameters per token, memory bandwidth, quantization). But with fixed hardware, reducing the number of input tokens — a shorter system prompt, pruned history — noticeably speeds up time-to-first-token.

#Count a text's tokens yourself

The best way to build intuition is to measure. Here are three approaches, from simplest to most precise.

#Estimate roughly

For a quick estimate without installing anything: divide the character count by 4 for English, or by 3 to 3.5 for French. It's rough but sufficient to determine whether a document fits in a context window.

#Count in Python with the actual tokenizer

For an exact count, use the model's tokenizer. Hugging Face's tokenizers / transformers library loads the actual tokenizer for an open-weight model:

Count with the model's tokenizer
from transformers import AutoTokenizer

# Remplacez par le modèle que vous faites tourner en local
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B")

texte_fr = "Le chat est sur la table."
texte_en = "The cat is on the table."

print(len(tok.encode(texte_fr)), "tokens (fr)")
print(len(tok.encode(texte_en)), "tokens (en)")

# Voir le découpage réel
print(tok.tokenize(texte_fr))

Running this script on your own text is the most revealing exercise: you can see for yourself which French words get split apart and how much more compact English is.

#Read the count from the Ollama API

Ollama already exposes token counts in its responses. The daemon listens by default on http://localhost:11434; an API call returns prompt_eval_count (prompt tokens) and eval_count (generated tokens):

Terminal
curl http://localhost:11434/api/generate -d '{
  "model": "qwen2.5:7b",
  "prompt": "Explique la tokenization en une phrase.",
  "stream": false
}' | grep -o '"eval_count":[0-9]*'

You’ll also find eval_duration: divide eval_count by the duration to get your actual throughput in tokens per second, on your machine, with your quantization.

#Save tokens without sacrificing quality

Since every token costs context and speed, a few simple habits make a real difference, especially in French.

Concise system prompt
It is sent with every call. Every unnecessary sentence costs you on every turn.
Trim history
Summarize or cut off old exchanges instead of dragging the entire conversation along.
RAG instead of pasting everything
Inject only the relevant passages from a document, not the entire document.
Targeting output length
Requesting a short response generates fewer tokens—and therefore responds faster.
→
The right order of magnitude
Before pasting a large document into the chat, estimate its size: ~3 characters per token in French. A text of 30 000 characters is ~10 000 tokens—one-third of a 32k window, even before your question and the response.

#Go further

Tokens are the thread connecting several basic concepts. Three guides build directly on this one: the first details the context window that tokens fill, the second explains how quantization affects speed measured in tokens per second, and the third shows how the transformer processes those tokens internally.


Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.