7B, 30B, 70B LLMs: what these sizes mean ?
On Hugging Face or Ollama, every model carries a number: 1B, 7B, 32B, 70B. This figure—the billions of parameters—determines what the model can do, how much VRAM it requires, and how fast it runs. But a recent 7B LLM can outperform a 70B model from two years ago, and “bigger” does not mean “better for you.” This guide explains what a parameter is, what each tier actually provides, and how to balance size, quantization, and hardware without being misled by marketing.
#What is a parameter, concretely?
A parameter is a number. Nothing more: a numerical value, often encoded in 16 bits, learned during training. A “7B” model contains 7 billion of them. These numbers are the weights of the connections between the network’s artificial neurons—they encode everything the model “knows”: grammar, facts, phrasing, and associations between ideas.
The most accurate analogy is that of synapses. A parameter is like the strength of a connection between two neurons in the brain. The more connections there are, the more nuances and subtle patterns the network can memorize. A 1-billion-parameter model has about 1 billion of these settings; a 70B model has seventy times as many. Each parameter is tiny and unintelligent in isolation; it is their number and arrangement that produce language.
In practice, these parameters take up space. At 16 bits (FP16), each parameter weighs 2 bytes, so a 7B model uses about 14 GB at full precision. This size on disk and in memory determines whether the model fits on your machine—we'll return to this with quantization, which compresses these weights.
#The real scale of capabilities, from 1B to 70B
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Capabilities do not increase linearly with size: they cross thresholds. Once a certain parameter threshold is passed, new abilities appear suddenly—this is what are called emergent capabilities. Here is what we observe in practice on recent models.
- 1B – 3B
- Text completion, simple rewriting, classification, and keyword extraction. Useful on mobile or at the edge, but reasoning is fragile and hallucinations are frequent as soon as the task becomes more complex.
- 7B – 8B
- The first genuinely useful tier. Coherent conversation, reliable summaries, simple code, and accurate instruction following. This is the entry point to serious local LLMs.
- 13B – 14B
- Better performance on long tasks, fewer factual errors, and stronger code. A notch more refined than the 7B, without blowing through your VRAM.
- 30B – 34B
- More reliable multistep reasoning, nuanced responses, and good code and translation quality. The quality/cost sweet spot for demanding local use.
- 70B
- The top of the consumer range. Cloud-model-like reasoning on many tasks, with detailed and nuanced answers. Requires a powerful setup or offloading.
Key takeaway: every doubling in size delivers a real but diminishing gain. The jump from 1B to 7B is spectacular; the jump from 30B to 70B is real but much subtler, and only makes sense if your task needs it.
#What a 7B LLM can do that a 1B cannot
This is the most telling comparison because the gap is enormous for an apparently modest difference in size. A 1B model repeats patterns; a 7B model starts to reason. Moving from one to the other unlocks concrete capabilities you notice immediately in use.
- Following a multi-part instruction
- A 1B model often forgets the second half of your request. A 7B model can handle an instruction like “summarize this text in 3 points, then translate them into English.”
- Chain-of-thought reasoning
- A logical or mathematical problem with two or three steps remains beyond a 1B model; a 7B model breaks it down correctly much of the time.
- Consistency across a long text
- The 1B loses the thread after a few paragraphs. The 7B retains the context, tone, and proper names throughout a complete response.
- Working code
- A specialized 7B model writes correct functions and fixes simple bugs; a 1B model mostly produces code that “looks” like code.
- Fewer hallucinations
- Without eliminating them, the 7B hallucinates significantly less and more often knows to say that it doesn't know.
#Size and VRAM: the table that matters
The practical question is not “which is the best model?” but “which model fits on my card?” The required VRAM directly tracks the parameter count. Here are the benchmarks for Q4_K_M quantization, the recommended default format, which reduces the size by approximately four compared with FP16.
- 3B ≈ 2 GB
- Runs anywhere, even on an entry-level card or CPU. Ideal for mobile and simple tasks.
- 7B ≈ 5 GB
- Comfortable on a RTX 3060 12 GB or a RTX 4070 12 GB. The best capacity-to-accessibility ratio.
- 14B ≈ 9 GB
- Fits on 12 GB of VRAM with a reasonable context, and is comfortable on 16 GB.
- 32B ≈ 19 GB
- Targets a RTX 4090 24 GB, or a 16 GB card with some offloading to RAM.
- 70B ≈ 40 GB
- Beyond the reach of a single consumer card: dual-GPU, RAM/SSD offload, or an Apple Apple Silicon Mac with generous unified memory (M4 Pro 48 GB).
#Size vs. quantization: the real trade-off
With limited VRAM, you have two options: choose a smaller model, or choose a larger but more compressed model. Quantization reduces parameter precision—from 16 bits to 4 or 8 bits—to fit the model into less memory, at the cost of a slight loss in quality.
The rule of thumb validated by experience: a larger, more heavily quantized model almost always beats a smaller full-precision model at the same VRAM capacity. A 13B model in Q4 fits in the same memory as a 7B model in Q8 and is generally more capable. Parameter count weighs more than the last few bits of precision.
- Q4_K_M
- The recommended default. Minimal quality loss, memory reduced by ~4. The default choice for nearly all use cases.
- Q5_K_M
- A step up in quality for a little more VRAM. A good compromise if memory allows.
- Q8_0
- Virtually indistinguishable from FP16. Best reserved for small models or very high-end cards.
- FP16
- Full precision, twice as large as Q8. Mainly useful for fine-tuning, rarely needed for inference.
#Why a small recent model beats a large old one
This is what confuses beginners most: a 7B LLM from 2026 can outperform a 70B model from 2023. Size is only one factor in quality—and not always the most decisive one. Three other levers have improved enormously.
- Data quality
- Recent models are trained on better-filtered, deduplicated corpora enriched with high-quality synthetic data. At the same size, they learn much more.
- Training volume
- Modern 7B models see tens of trillions of tokens, far more than older large models. A small network trained extensively outperforms a large undertrained network.
- Architecture and post-training
- Better architectures, RLHF alignment, and reasoning-specific training—all gains that have nothing to do with parameter count.
Distillation also plays a key role: knowledge is transferred from a very large model to a small one, which inherits some of its capabilities at a fraction of the size. That is how a distilled 7B or 8B can reason like a much larger model. Practical consequence: always check the release date and recent benchmarks, never just the parameter count.
#Choosing the right size in practice
The right size is the one that fits your hardware while covering your use case. Start with the largest recent model your VRAM can handle in Q4, then adjust based on your actual need for speed or quality.
- 01Measure your VRAMIdentify your GPU memory (or unified memory on Mac). This is the hard ceiling that immediately rules out models that are too large.
- 02Aim for the largest recent model that fits in Q4On 12 GB, a recent 14B. On 24 GB, a 32B. On 8 GB, a good 7B/8B. Always prefer a recent release over an older, larger one.
- 03Adjust based on usageEveryday chat and speed: stick with 7B–14B. Reasoning, demanding code, fine-grained translation: move up to 30B+ if memory allows.
- 04Check the speedIf the tokens-per-second throughput frustrates you, move down a tier. A fast 7B often beats a sluggish 32B for interactive use.
#Compare two sizes in 5 minutes
The best way to feel the difference is to run two sizes side by side on your own machine. If Ollama is installed and listening on its default port (http://localhost:11434), two commands are enough.
Ask both the same reasoning question (for example, a small multistep logic problem) and compare the answer quality and the throughput shown by --verbose. You’ll see exactly where the capacity/speed tradeoff lies on your hardware.
#Go further
Size is easier to understand when connected to the other fundamentals of local models. These guides build directly on this one:
- Choose your quantization (Q4, Q5, Q8, FP16)
- The other half of the memory tradeoff: how to compress a model without degrading it, with a visual guide.
- MoE explained: why a 30B-A3B runs like a small model
- The two-digit notation that breaks the link between size and speed, to read right after this one.
- Distillation: how small models inherit from large ones
- To understand in depth why a small, recent model can beat a large, older one.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.