Beginner 12 minBasics

An LLM's architecture: the transformer explained simplement

People talk everywhere about “transformers,” “attention,” and “parameters” without ever saying what these words actually mean. This guide opens the hood on an LLM and explains its architecture with simple analogies, without a single equation. By the end, you will understand llm architecture from the inside—and, above all, why these design choices determine the VRAM the model requires and the speed at which it responds on your machine.

By Mohamed Meguedmi·Update 2026-09-12·Tested on Windows, macOS, and Linux

#Why understanding an LLM’s architecture matters

You can run an LLM locally without knowing anything about how it works internally—a ollama run suffit. But as soon as you want to choose the right model for your machine, every term in the technical specifications becomes an obstacle: “32 layers,” “7 billion parameters,” “MoE 8x7B,” “32-head attention.” These are the figures that determine whether a model will fit on your graphics card or crawl on your processor.

The good news: the architecture that dominates all current LLMs—the transformer—rests on a handful of ideas that can be explained without mathematics. Understanding these ideas takes you from “I copy a command without knowing why” to “I know why this model needs 9 GB of VRAM, not 40.”

i
No formulas in this guide
Everything is explained through analogies. If you’re looking for the mathematical details (dot product, softmax, positional encoding), this guide isn’t for that—it aims for the right intuition, enough to choose and run a model locally.

#Transforming it into an image

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A modern LLM is a transformer: a machine that takes text as input and predicts the next word, over and over. The name comes from the paper “Attention Is All You Need” (Google, 2017), which introduced this architecture. All the models you encounter — Llama, Qwen, Mistral, Gemma, DeepSeek, Phi — are variants of it.

Imagine an assembly line. At the input is your sentence, split into pieces. Each station on the line (a “layer”) deepens its understanding of the text while taking the context into account. At the output, the model proposes the most likely next piece. The process repeats for each new piece generated. That's it — everything else is just detail about what each station does.

Entry
Your text, split into tokens (small word fragments).
Layer stacking
Each layer refines the representation of the text. A model has dozens of them.
Warning
The core mechanism of every layer: it makes each word “look” at the others.
Output
A probability for each possible token; the model chooses one, adds it to the text, and repeats.

#From your words to tokens

An LLM does not see letters or even whole words: it sees tokens. A token is a frequent text fragment—sometimes a complete short word (“chat”), sometimes part of a word (“anti”, “constitution”), and sometimes a space or punctuation. In French, roughly count on 3 tokens for 2 words.

Each token is then converted into a list of numbers called an “embedding.” This translates the text into a language the machine can work with: coordinates in a space where words with similar meanings are geographically close. “King” and “queen” are neighbors there; “king” and “broccoli” are far apart.

→
The link to the context window
A model’s “context window” (for example, 8k, 32k, 128k) is measured in tokens, not words. A context of 8,000 tokens ≈ 6,000 French words, or about ten pages. This is the amount of text the model can “keep in mind” at once.

#Attention, the heart of the transformer

Attention is the idea that changed everything. Consider the sentence “The mouse ate the cheese because it was hungry.” What does “it” refer to? The mouse, obviously. To determine that, you need to connect “it” to “mouse,” several words earlier. That's exactly what attention does: for each word, it decides which other words in the sentence are relevant and how much to consider them.

Analogy: in a meeting, when you hear an ambiguous pronoun, your brain scans what was said earlier to figure out who is being referred to. Attention does the same thing, in parallel, for all words at once. Each word “asks a question” (what do I need to be understood?) and “receives answers” from the other words, weighted by their relevance.

People often talk about “attention heads.” A head is one way of looking at relationships; having several (32 heads, 64 heads…) lets the model track multiple types of connections at once—grammar on one side, the subject of the text on another, and temporal references elsewhere.

i
Why attention is expensive in memory
Every word looks at every other word. The longer the context, the more the number of relationships to store (the “KV cache”) grows explosively. That is why a very large context (128k tokens) can consume as much VRAM as the model weights themselves.

#Stacked layers

A single attention layer understands simple relationships. The power comes from stacking them: the output of one layer becomes the input to the next. A small model has around twenty; a large model has several dozen. Each layer combines two blocks: attention (which connects words to one another) and a “feed-forward” network (which processes each word individually, like a mini-brain thinking about what it just read).

Analogy: reading at multiple levels. The first layer identifies words and grammar. The middle layers construct the meaning of sentences. The last layers capture the intent, tone, and what should come next. The more layers there are, the more deeply the model can reason—but the more computation is required for each generated token.

Attention block
Connects words to one another (the context).
Feed-forward block
Transforms each word in depth (the model's “knowledge”).
Number of layers
“Depth.” The more there is, the more capable—and slower—the model is.
Width (hidden dimension)
The size of the number lists. The larger it is, the “larger” and heavier the model.

#What exactly are parameters?

When you see “7B” or “70B,” the B means “billion” and refers to the number of parameters. A parameter is a tunable number inside the model—one of countless knobs adjusted during training. These knobs encode everything the model “knows”: grammar, facts, styles, and reasoning.

Analogy: imagine a giant mixing console with billions of sliders. Training consists of adjusting each slider so that, across billions of text examples, the model predicts the correct next word. Once fixed, these settings are the “weights” that you download when you make a ollama pull.

The more parameters a model has, the more it can memorize and distinguish—but the more memory it uses and the slower it is, because every generated token passes through all those parameters. This is the direct link between the “7B” figure and your machine’s hardware requirements.

→
Parameters ≠ bytes
A parameter does not occupy 1 byte. At native precision (FP16), it takes 2. A 7B model therefore weighs ~14 GB in FP16. Quantization reduces this size by storing each parameter with fewer bits—that is where Q4, Q5, and Q8 come in (see below).

#Dense vs. MoE: two ways to wire the model

So far, we have described a “dense” transformer: every token passes through all the parameters. Simple, but expensive—a dense 70B calls on all 70 billion parameters for every generated word.

The MoE (Mixture of Experts) architecture breaks this rule. Instead of a single large feed-forward block per layer, it uses several (the “experts”) and activates only a few for each token, selected by a small dispatcher (the “router”). Analogy: rather than a generalist who answers everything, a practice of specialists where you consult only the two doctors relevant to your case.

Dense
All parameters are active for every token. E.g.: Llama 3 8B, Qwen 14B, Gemma 27B.
MoE
Many parameters overall, but few active per token. E.g.: Mixtral 8x7B, DeepSeek V3, Llama 4 Scout.
MoE notation
“30B-A3B” = 30 billion parameters in total, but only 3 billion active parameters (A = active) per token.

The consequence for local setups is major: an MoE behaves like a small model in terms of speed (few active parameters) while retaining the knowledge of a large model (many total parameters). The catch is VRAM: you must load all the experts into memory, even though only a fraction are activated at a time.


#Why all of this determines VRAM and speed

Here’s the practical core. Two resources matter locally: memory (to hold the model) and compute (to respond quickly). The architecture governs both.

#Memory: the weights have to fit

To be fast, a model must fit entirely in your GPU's VRAM (or the unified memory of a Mac Apple Silicon). Otherwise, part of it spills over to the processor and RAM, and speed collapses. The size depends on the number of parameters and the quantization.

3B in Q4
≈ 2 GB of VRAM
7B in Q4
≈ 5 GB
14B in Q4
≈ 9 GB
32B in Q4
≈ 19 GB
70B in Q4
≈ 40 GB

Add the context memory (the KV cache), which grows with prompt length. A long context can require several additional gigabytes—don’t forget this when you’re right at your card’s limit.

#Speed: how many active parameters per token

Generation speed (tokens per second) depends mainly on the parameters actually activated for each token. That is why a dense 8B and a 30B-A3B MoE can have comparable speed: both activate only ~3 to 8 billion parameters per token. The MoE simply requires much more VRAM to hold all its experts.

!
The classic overflow trap
A model that “almost fits” in your VRAM does not fit. As soon as some layers are offloaded to the CPU, you lose a factor of 5 to 20 in speed. It's better to use a model one step smaller (or a more aggressive quantization) that fits 100% on the GPU.
RTX 3060 12 GB
Comfortable up to a 14B in Q4. Ideal entry-level card.
RTX 4070 / 4080 (12–16 GB)
14B comfortably, 32B in Q4 tight on the 4080.
RTX 4090 24 GB
32B is comfortable in Q4; 70B is out of reach on pure GPU.
Mac M4 Pro 24–48 GB unified memory
Unified memory serves as VRAM: up to a 70B model in Q4 on 48 GB configurations.

#Read a model card line by line

You now have the keys to decode a technical specification sheet. Let’s take a typical example, like those found on Hugging Face or the Ollama library:

Model card (typical excerpt)
{
  "architecture": "transformer (decoder-only)",
  "parameters": "14B",
  "type": "dense",
  "layers": 40,
  "attention_heads": 40,
  "context_length": 32768,
  "quantization": "Q4_K_M",
  "size_on_disk": "9 GB"
}
architecture: decoder-only
The standard for generative LLMs: the model only predicts the next part of the text (with no separate “encoder” component).
parameters: 14B
14 billion parameters. Memory reference: ~9 GB in Q4; at least a 12 GB card is required.
type: dense
All parameters are active at every token. If it were an MoE, you would see notation such as 30B-A3B.
layers: 40
40 stacked layers. That's depth: the more layers there are, the more nuanced the reasoning, and the slower it is.
attention_heads: 40
40 attention heads per layer—just as many ways to connect words in parallel.
context_length: 32768
32k context tokens, or ~24,000 words. Note: filling this context consumes additional VRAM.
quantization: Q4_K_M
Reduced to ~4 bits per parameter. The best quality-to-size tradeoff for local use.
size_on_disk: 9 GB
What you download and what needs to fit in memory for maximum speed.

With these seven lines, you can already answer the only question that matters: “does it run well on my machine?” Here: 14B in Q4 = ~9 GB of weights + context → a RTX 3060 12 GB or better, and it runs fast.

→
The rule of thumb to remember
Two figures are enough for an initial screening: the parameter count (and whether it is dense or MoE) for VRAM, and the quantization for the actual size. The rest (layers, heads, context) refines the assessment but does not change basic feasibility.

#Go further

You now know the anatomy of an LLM. Three guides naturally build on these foundations: the first explains MoE notation and its practical impact, the second explains the choice of quantization that determines on-disk size, and the third revisits the context window discussed here.


Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.