Beginner 8 minConcepts

7B, 30B, 70B LLMs: what these sizes mean ?

On Hugging Face or Ollama, every model carries a number: 1B, 7B, 32B, 70B. This figure—the billions of parameters—determines what the model can do, how much VRAM it requires, and how fast it runs. But a recent 7B LLM can outperform a 70B model from two years ago, and “bigger” does not mean “better for you.” This guide explains what a parameter is, what each tier actually provides, and how to balance size, quantization, and hardware without being misled by marketing.

By Clara M.·Update 2026-09-21·Tested on Windows, macOS, and Linux

#What is a parameter, concretely?

A parameter is a number. Nothing more: a numerical value, often encoded in 16 bits, learned during training. A “7B” model contains 7 billion of them. These numbers are the weights of the connections between the network’s artificial neurons—they encode everything the model “knows”: grammar, facts, phrasing, and associations between ideas.

The most accurate analogy is that of synapses. A parameter is like the strength of a connection between two neurons in the brain. The more connections there are, the more nuances and subtle patterns the network can memorize. A 1-billion-parameter model has about 1 billion of these settings; a 70B model has seventy times as many. Each parameter is tiny and unintelligent in isolation; it is their number and arrangement that produce language.

i
Why the “B”
The B means billion in English, or one milliard in French. 7B = 7 billion parameters. This figure says nothing about the quality of the training data or the architecture—only the network’s raw size.

In practice, these parameters take up space. At 16 bits (FP16), each parameter weighs 2 bytes, so a 7B model uses about 14 GB at full precision. This size on disk and in memory determines whether the model fits on your machine—we'll return to this with quantization, which compresses these weights.

#The real scale of capabilities, from 1B to 70B

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Capabilities do not increase linearly with size: they cross thresholds. Once a certain parameter threshold is passed, new abilities appear suddenly—this is what are called emergent capabilities. Here is what we observe in practice on recent models.

1B – 3B
Text completion, simple rewriting, classification, and keyword extraction. Useful on mobile or at the edge, but reasoning is fragile and hallucinations are frequent as soon as the task becomes more complex.
7B – 8B
The first genuinely useful tier. Coherent conversation, reliable summaries, simple code, and accurate instruction following. This is the entry point to serious local LLMs.
13B – 14B
Better performance on long tasks, fewer factual errors, and stronger code. A notch more refined than the 7B, without blowing through your VRAM.
30B – 34B
More reliable multistep reasoning, nuanced responses, and good code and translation quality. The quality/cost sweet spot for demanding local use.
70B
The top of the consumer range. Cloud-model-like reasoning on many tasks, with detailed and nuanced answers. Requires a powerful setup or offloading.

Key takeaway: every doubling in size delivers a real but diminishing gain. The jump from 1B to 7B is spectacular; the jump from 30B to 70B is real but much subtler, and only makes sense if your task needs it.

#What a 7B LLM can do that a 1B cannot

This is the most telling comparison because the gap is enormous for an apparently modest difference in size. A 1B model repeats patterns; a 7B model starts to reason. Moving from one to the other unlocks concrete capabilities you notice immediately in use.

Following a multi-part instruction
A 1B model often forgets the second half of your request. A 7B model can handle an instruction like “summarize this text in 3 points, then translate them into English.”
Chain-of-thought reasoning
A logical or mathematical problem with two or three steps remains beyond a 1B model; a 7B model breaks it down correctly much of the time.
Consistency across a long text
The 1B loses the thread after a few paragraphs. The 7B retains the context, tone, and proper names throughout a complete response.
Working code
A specialized 7B model writes correct functions and fixes simple bugs; a 1B model mostly produces code that “looks” like code.
Fewer hallucinations
Without eliminating them, the 7B hallucinates significantly less and more often knows to say that it doesn't know.
→
The 7B threshold
If your machine allows it, always start with a 7B/8B rather than a 3B. The reliability gain is disproportionate to the extra VRAM cost (about 5 GB in Q4).

#Size and VRAM: the table that matters

The practical question is not “which is the best model?” but “which model fits on my card?” The required VRAM directly tracks the parameter count. Here are the benchmarks for Q4_K_M quantization, the recommended default format, which reduces the size by approximately four compared with FP16.

3B ≈ 2 GB
Runs anywhere, even on an entry-level card or CPU. Ideal for mobile and simple tasks.
7B ≈ 5 GB
Comfortable on a RTX 3060 12 GB or a RTX 4070 12 GB. The best capacity-to-accessibility ratio.
14B ≈ 9 GB
Fits on 12 GB of VRAM with a reasonable context, and is comfortable on 16 GB.
32B ≈ 19 GB
Targets a RTX 4090 24 GB, or a 16 GB card with some offloading to RAM.
70B ≈ 40 GB
Beyond the reach of a single consumer card: dual-GPU, RAM/SSD offload, or an Apple Apple Silicon Mac with generous unified memory (M4 Pro 48 GB).
!
Don't forget the context
These figures cover the model weights alone. The context window (the KV cache) consumes additional VRAM, especially when it is large. Allow a margin of 1 to a few GB above the model size.

#Size vs. quantization: the real trade-off

With limited VRAM, you have two options: choose a smaller model, or choose a larger but more compressed model. Quantization reduces parameter precision—from 16 bits to 4 or 8 bits—to fit the model into less memory, at the cost of a slight loss in quality.

The rule of thumb validated by experience: a larger, more heavily quantized model almost always beats a smaller full-precision model at the same VRAM capacity. A 13B model in Q4 fits in the same memory as a 7B model in Q8 and is generally more capable. Parameter count weighs more than the last few bits of precision.

Q4_K_M
The recommended default. Minimal quality loss, memory reduced by ~4. The default choice for nearly all use cases.
Q5_K_M
A step up in quality for a little more VRAM. A good compromise if memory allows.
Q8_0
Virtually indistinguishable from FP16. Best reserved for small models or very high-end cards.
FP16
Full precision, twice as large as Q8. Mainly useful for fine-tuning, rarely needed for inference.
→
The trade-off in one sentence
With the same memory capacity, prefer a larger model in Q4 over a smaller model in Q8. Do not go below Q4 unless necessary: below that, quality drops quickly.

#Why a small recent model beats a large old one

This is what confuses beginners most: a 7B LLM from 2026 can outperform a 70B model from 2023. Size is only one factor in quality—and not always the most decisive one. Three other levers have improved enormously.

Data quality
Recent models are trained on better-filtered, deduplicated corpora enriched with high-quality synthetic data. At the same size, they learn much more.
Training volume
Modern 7B models see tens of trillions of tokens, far more than older large models. A small network trained extensively outperforms a large undertrained network.
Architecture and post-training
Better architectures, RLHF alignment, and reasoning-specific training—all gains that have nothing to do with parameter count.

Distillation also plays a key role: knowledge is transferred from a very large model to a small one, which inherits some of its capabilities at a fraction of the size. That is how a distilled 7B or 8B can reason like a much larger model. Practical consequence: always check the release date and recent benchmarks, never just the parameter count.

!
The marketing trap
“Bigger = better” is not universally true. A model is good because of its data, training, and architecture as much as its size. Beware of comparisons that line up only the billions of parameters.

#Choosing the right size in practice

The right size is the one that fits your hardware while covering your use case. Start with the largest recent model your VRAM can handle in Q4, then adjust based on your actual need for speed or quality.

  1. 01
    Measure your VRAM
    Identify your GPU memory (or unified memory on Mac). This is the hard ceiling that immediately rules out models that are too large.
  2. 02
    Aim for the largest recent model that fits in Q4
    On 12 GB, a recent 14B. On 24 GB, a 32B. On 8 GB, a good 7B/8B. Always prefer a recent release over an older, larger one.
  3. 03
    Adjust based on usage
    Everyday chat and speed: stick with 7B–14B. Reasoning, demanding code, fine-grained translation: move up to 30B+ if memory allows.
  4. 04
    Check the speed
    If the tokens-per-second throughput frustrates you, move down a tier. A fast 7B often beats a sluggish 32B for interactive use.

#Compare two sizes in 5 minutes

The best way to feel the difference is to run two sizes side by side on your own machine. If Ollama is installed and listening on its default port (http://localhost:11434), two commands are enough.

Terminal
# Un petit modèle rapide
ollama run llama3.2:3b --verbose

# Le même prompt sur un modèle plus gros
ollama run llama3.1:8b --verbose

Ask both the same reasoning question (for example, a small multistep logic problem) and compare the answer quality and the throughput shown by --verbose. You’ll see exactly where the capacity/speed tradeoff lies on your hardware.

i
Read tokens per second
The --verbose flag displays the number of tokens per second at the end of the response. This is the objective measurement for choosing between two sizes on your own GPU.

#Go further

Size is easier to understand when connected to the other fundamentals of local models. These guides build directly on this one:

Choose your quantization (Q4, Q5, Q8, FP16)
The other half of the memory tradeoff: how to compress a model without degrading it, with a visual guide.
MoE explained: why a 30B-A3B runs like a small model
The two-digit notation that breaks the link between size and speed, to read right after this one.
Distillation: how small models inherit from large ones
To understand in depth why a small, recent model can beat a large, older one.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.