LLM benchmarks: understanding leaderboards (MMLU, Arena, SWE-bench)
Every new model comes with a score chart placing it “at GPT-5 level.” But an LLM benchmark never measures “intelligence”: it measures a specific task, with a specific protocol, and often a specific weakness. This guide surveys the major leaderboards (MMLU, GPQA, LMArena, SWE-bench), explains their pitfalls—contamination, saturation, cherry-picking—and shows how to evaluate a local model yourself on what really matters: your tasks.
#Why benchmarks matter (and mislead you)
An LLM benchmark is a set of questions with known answers that you ask a model to count how many it gets right. The result is a score—a percentage, an Elo ranking, or a resolution rate. It is the only common language that lets you compare two models without testing them yourself for hours, which is why everyone uses it.
The problem is not the principle but the gap between what the score says and what you read into it. A model that scores 90% on MMLU is not “90% intelligent”: it answers 90% of an academic general-knowledge multiple-choice test correctly. That says nothing about its ability to follow your instructions, write correct French, avoid hallucinating in your domain, or sustain a ten-turn conversation. A benchmark measures a narrow skill under laboratory conditions.
Benchmarks can be grouped into three broad families, which do not measure the same thing and are not used in the same way: academic benchmarks (knowledge and reasoning multiple-choice tests), real-task benchmarks (solving a real bug, using tools), and human-preference arenas (humans vote for the best answer). We'll take them in order.
#Academic benchmarks: MMLU, GPQA, MATH
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
This is the traditional family, the one behind the score tables seen on model release pages. They are sets of questions with verifiable answers, often multiple-choice, scored automatically. Their strength is reproducibility; their weakness is that they are easy to saturate and contaminate.
- MMLU
- Massive Multitask Language Understanding: ~14,000 multiple-choice questions across 57 subjects (law, medicine, history, math…). The most widely cited knowledge benchmark. Now largely saturated: strong models exceed 88–90%, and the gap between them is mostly noise.
- MMLU-Pro
- A hardened version of MMLU: 10 choices instead of 4, questions reworked to be more difficult, and more reasoning. Created specifically because MMLU no longer distinguished the leading models.
- GPQA (Diamond)
- “Google-Proof Q&A”: doctoral-level questions in biology, physics, and chemistry designed so that a Google search alone isn’t enough. The “Diamond” subset is the hardest. A good indicator of cutting-edge scientific reasoning.
- MATH / AIME
- Math problems (MATH: high-school/competition level; AIME: American olympiads). A single numerical answer, so they are easy to grade. It has become a benchmark for “reasoning” models that think through multiple steps.
- IFEval
- Measures the ability to FOLLOW verifiable instructions (“respond in exactly 3 bullet points,” “don’t use the letter e”). Often more revealing for real-world use than a knowledge quiz.
- HLE (Humanity's Last Exam)
- Recent, deliberately extreme benchmark: expert-level, multi-domain questions designed to remain difficult for a long time. The best models still score low on it, making it a good discriminator in 2026.
#Real-world task benchmarks: SWE-bench and agents
This family is newer and much harder to game because it does not ask questions: it asks you to complete an entire task whose result can be verified objectively. Resolve a real GitHub issue, get a test suite to pass, or navigate a code repository. We cover this in detail in our dedicated guide to code benchmarks; here is the essential information.
- SWE-bench Verified
- A manually validated subset of 500 real GitHub problems (issues + expected patch) that OpenAI deemed solvable and well specified. The model must produce a patch that makes the repository's tests pass. It has become THE standard for real-world coding, far more informative than HumanEval (now saturated at 96-98%).
- SWE-bench (full / Lite)
- The full version (~2,300 tasks) and a streamlined version for fast iteration. Published scores depend enormously on the “harness” (the agent that orchestrates the model): the same model can gain 15 points depending on the surrounding tooling.
- Tau-bench / agentic tasks
- Agent benchmarks that use tools (function calls, APIs, compliance with business rules) over multiple turns. They measure reliability in “assistant that acts” use, not just one that answers.
- LiveCodeBench
- Continuously collected competitive programming problems, with a publication date. You can score only problems published after a model's training date—an immediate antidote to contamination.
#LMArena: ranking by human preference
LMArena (formerly LMSYS's “Chatbot Arena”) works differently: two anonymous models answer the same question submitted by a real user, who votes for the better answer. Based on millions of head-to-head matches, an Elo score is calculated (the same system used in chess). It is the benchmark closest to “which model do people actually prefer?”
- What it measures well
- Perceived quality in open-ended conversation: tone, structure, overall usefulness, and ability to provide a pleasant response. It’s strongly correlated with satisfaction in real-world use.
- What it measures poorly
- Factual accuracy. Voters often reward long, well-formatted, confident answers—even when they’re wrong. A “flattering” model can rise in the rankings without being more accurate.
- Elo is not a percentage
- A 10–20 Elo-point difference is statistical noise. Always look at the displayed confidence interval: two models can be “tied” even if they are not on the same row of the table.
- Sub-rankings
- LMArena offers categories (code, math, long-form answers, controlled style). The “hard prompts” or “style control” sub-ranking is often more informative than the overall ranking, which mixes everything together.
#The #1 pitfall: data contamination
Contamination is when benchmark questions—or their answers—end up, intentionally or otherwise, in the model's training data. The model is no longer reasoning: it is reciting. Public benchmarks circulate on the web, GitHub, and Hugging Face—they inevitably end up in pretraining corpora. The result is an inflated score that predicts nothing about unseen questions.
This is the Achilles' heel of all static, public benchmarks. The older and more famous a benchmark is, the higher the risk. Warning signs include:
- Differences between versions
- A model that crushes MMLU but collapses on MMLU-Pro (same subjects, new questions) looks more like memorization than understanding.
- Abnormally high score for its size
- A small 7B model that beats 70B models on one specific benchmark and only that one: be wary—it is often targeted training (“benchmaxxing”) on that test set.
- Dated benchmarks
- “Live” evaluations (LiveCodeBench, time-stamped questions) work around the problem: only material after the model's cutoff date is scored.
- Suspected memory leaks
- Some tests inject canary variants (“canary strings”) or paraphrases to detect memorization. A large gap between the original and paraphrased question reveals contamination.
#Pitfall No. 2: saturation
A benchmark is saturated when the best models reach scores so high that it can no longer distinguish them. When everyone scores 96-99%, the remaining 3 points are measurement noise (ambiguous questions, labeling errors) rather than a real difference in capability. HumanEval (code) and MMLU (knowledge) are the canonical examples: they were very useful, but are no longer useful for separating the top of the rankings.
That is why tougher benchmarks keep appearing: MMLU-Pro replaces MMLU, GPQA Diamond and HLE take over for reasoning, and SWE-bench Verified replaces HumanEval for code. A benchmark has a useful lifespan; once the average level passes a certain point, remove it from your decision criteria.
#How to read a leaderboard without being misled
Here is the method to apply to any score table, whether it comes from a model announcement or a public ranking.
- 01Identify who publishes itA chart in a model announcement is marketing: it selects favorable benchmarks and an advantageous protocol (cherry-picking). A neutral, third-party leaderboard (Open LLM Leaderboard, LMArena, the official SWE-bench page) is far more reliable than a launch slide.
- 02Check the protocol0-shot or few-shot? With chain-of-thought? Which harness for agentic systems? Two scores can only be compared if the method is identical. A “self-reported” asterisk is worth less than a score reproduced by a third party.
- 03Look at the confidence intervalsOn LMArena, an Elo gap of less than ~15 points is noise. On low-sample benchmarks (GPQA Diamond, 198 questions), a handful of correct answers can move the score by several points. A ranking without a margin of error should be taken with a grain of salt.
- 04Cross-reference multiple benchmarksNever decide based on a single number. A solid model performs well across a varied family of tests, not just one peak. An isolated, abnormal peak is more suggestive of targeted optimization.
- 05Weight it according to YOUR use caseDo you write code? Look at SWE-bench and LiveCodeBench, not MMLU. French? None of these benchmarks are in French—look for French evaluations or test them yourself. Conversation? LMArena is the priority. The best model “on average” is not necessarily the best one for you.
#Evaluate a local model on your own tasks
The logical conclusion from everything above: the most reliable benchmark for you is your own. It cannot be contaminated (your questions are not on the web), it is never saturated (you calibrate it to your difficult cases), and it measures exactly what you need. No heavy infrastructure is required: around twenty representative examples are enough to distinguish between two models.
- 01Build a mini test setGather 15 to 30 prompts from your real use cases (emails to draft, extractions, questions about your documents, code snippets). For each one, note what a good answer must contain. Keep this set private and stable so you can compare models over time.
- 02Install the models to compareWith Ollama, pull the candidate models. Ollama exposes an OpenAI-compatible API on http://localhost:11434, allowing calls to be automated.
- 03Automate callsA small script loops through your prompts and records each model's responses side by side in a file, so you can compare them later without being influenced by the model name.
- 04Record results blindlyReview the answers without knowing which model produced what (shuffle the order). Score them using your own criteria: accuracy, adherence to the instructions, quality of the French, and absence of fabrication. This is your personal “Arena.”
- 05Also measure the real costQuality is not everything: track speed (tokens per second) and VRAM usage. A 14B model in Q4_K_M (~9 GB) that fits on your RTX 4070 and responds quickly may outperform a 70B model (~40 GB) that is theoretically “better” but unusable on your system.
#Go further
These guides build on the concepts covered here, focusing on specialized benchmarks and the practical choice of a local model:
- Code benchmarks in detail
- “HumanEval Is Dead: Understanding LLM Code Benchmarks in 2026” explores SWE-bench, LiveCodeBench, and how to read code scores—the direct development-focused continuation of this guide.
- Choosing your quantization
- “GGUF Quantization in 2026: Q4_K_M vs Q5_K_M vs Q6_K” shows how to measure the real impact of quantization on quality—an in-house evaluation applied to a concrete case.
- A complete model test
- “Qwen 3 locally: complete test and real-world benchmarks” illustrates the local evaluation method (tokens/sec, French quality, VRAM) on a specific model.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.