Intermediate 11 minBenchmarks

LLM benchmarks: understanding leaderboards (MMLU, Arena, SWE-bench)

Every new model comes with a score chart placing it “at GPT-5 level.” But an LLM benchmark never measures “intelligence”: it measures a specific task, with a specific protocol, and often a specific weakness. This guide surveys the major leaderboards (MMLU, GPQA, LMArena, SWE-bench), explains their pitfalls—contamination, saturation, cherry-picking—and shows how to evaluate a local model yourself on what really matters: your tasks.

By Mohamed Meguedmi·Update 2026-09-04·Tested on Windows, macOS, and Linux

#Why benchmarks matter (and mislead you)

An LLM benchmark is a set of questions with known answers that you ask a model to count how many it gets right. The result is a score—a percentage, an Elo ranking, or a resolution rate. It is the only common language that lets you compare two models without testing them yourself for hours, which is why everyone uses it.

The problem is not the principle but the gap between what the score says and what you read into it. A model that scores 90% on MMLU is not “90% intelligent”: it answers 90% of an academic general-knowledge multiple-choice test correctly. That says nothing about its ability to follow your instructions, write correct French, avoid hallucinating in your domain, or sustain a ten-turn conversation. A benchmark measures a narrow skill under laboratory conditions.

i
The rule to keep in mind
A benchmark answers the question “does this model succeed at THIS specific task?”—never “is this model better for MY use case?”. The two sometimes overlap, often only partially, and never completely.

Benchmarks can be grouped into three broad families, which do not measure the same thing and are not used in the same way: academic benchmarks (knowledge and reasoning multiple-choice tests), real-task benchmarks (solving a real bug, using tools), and human-preference arenas (humans vote for the best answer). We'll take them in order.

#Academic benchmarks: MMLU, GPQA, MATH

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

This is the traditional family, the one behind the score tables seen on model release pages. They are sets of questions with verifiable answers, often multiple-choice, scored automatically. Their strength is reproducibility; their weakness is that they are easy to saturate and contaminate.

MMLU
Massive Multitask Language Understanding: ~14,000 multiple-choice questions across 57 subjects (law, medicine, history, math…). The most widely cited knowledge benchmark. Now largely saturated: strong models exceed 88–90%, and the gap between them is mostly noise.
MMLU-Pro
A hardened version of MMLU: 10 choices instead of 4, questions reworked to be more difficult, and more reasoning. Created specifically because MMLU no longer distinguished the leading models.
GPQA (Diamond)
“Google-Proof Q&A”: doctoral-level questions in biology, physics, and chemistry designed so that a Google search alone isn’t enough. The “Diamond” subset is the hardest. A good indicator of cutting-edge scientific reasoning.
MATH / AIME
Math problems (MATH: high-school/competition level; AIME: American olympiads). A single numerical answer, so they are easy to grade. It has become a benchmark for “reasoning” models that think through multiple steps.
IFEval
Measures the ability to FOLLOW verifiable instructions (“respond in exactly 3 bullet points,” “don’t use the letter e”). Often more revealing for real-world use than a knowledge quiz.
HLE (Humanity's Last Exam)
Recent, deliberately extreme benchmark: expert-level, multi-domain questions designed to remain difficult for a long time. The best models still score low on it, making it a good discriminator in 2026.
!
Watch the measurement protocol
The same model can receive two different MMLU scores depending on the method: 0-shot versus 5-shot (with examples), with or without chain-of-thought, using an exact prompt versus a rephrased one. Comparing two scores measured under different conditions is meaningless. Always check that the comparison is made “under identical conditions.”

#Real-world task benchmarks: SWE-bench and agents

This family is newer and much harder to game because it does not ask questions: it asks you to complete an entire task whose result can be verified objectively. Resolve a real GitHub issue, get a test suite to pass, or navigate a code repository. We cover this in detail in our dedicated guide to code benchmarks; here is the essential information.

SWE-bench Verified
A manually validated subset of 500 real GitHub problems (issues + expected patch) that OpenAI deemed solvable and well specified. The model must produce a patch that makes the repository's tests pass. It has become THE standard for real-world coding, far more informative than HumanEval (now saturated at 96-98%).
SWE-bench (full / Lite)
The full version (~2,300 tasks) and a streamlined version for fast iteration. Published scores depend enormously on the “harness” (the agent that orchestrates the model): the same model can gain 15 points depending on the surrounding tooling.
Tau-bench / agentic tasks
Agent benchmarks that use tools (function calls, APIs, compliance with business rules) over multiple turns. They measure reliability in “assistant that acts” use, not just one that answers.
LiveCodeBench
Continuously collected competitive programming problems, with a publication date. You can score only problems published after a model's training date—an immediate antidote to contamination.
i
Why these benchmarks are more reliable
A patch that makes the tests pass is objectively verifiable and difficult to memorize blindly: the model must actually produce a working solution. That's why agentic benchmarks resist saturation better than multiple-choice tests.

#LMArena: ranking by human preference

LMArena (formerly LMSYS's “Chatbot Arena”) works differently: two anonymous models answer the same question submitted by a real user, who votes for the better answer. Based on millions of head-to-head matches, an Elo score is calculated (the same system used in chess). It is the benchmark closest to “which model do people actually prefer?”

What it measures well
Perceived quality in open-ended conversation: tone, structure, overall usefulness, and ability to provide a pleasant response. It’s strongly correlated with satisfaction in real-world use.
What it measures poorly
Factual accuracy. Voters often reward long, well-formatted, confident answers—even when they’re wrong. A “flattering” model can rise in the rankings without being more accurate.
Elo is not a percentage
A 10–20 Elo-point difference is statistical noise. Always look at the displayed confidence interval: two models can be “tied” even if they are not on the same row of the table.
Sub-rankings
LMArena offers categories (code, math, long-form answers, controlled style). The “hard prompts” or “style control” sub-ranking is often more informative than the overall ranking, which mixes everything together.
→
Bridge the two worlds
A model that is strong on academic benchmarks but weak in Arena is often a “good student but rigid.” A model that is strong in Arena but average on GPQA is often “pleasant but less reliable in substance.” The right compromise lies at the intersection of the two, not in a single ranking.

#The #1 pitfall: data contamination

Contamination is when benchmark questions—or their answers—end up, intentionally or otherwise, in the model's training data. The model is no longer reasoning: it is reciting. Public benchmarks circulate on the web, GitHub, and Hugging Face—they inevitably end up in pretraining corpora. The result is an inflated score that predicts nothing about unseen questions.

This is the Achilles' heel of all static, public benchmarks. The older and more famous a benchmark is, the higher the risk. Warning signs include:

Differences between versions
A model that crushes MMLU but collapses on MMLU-Pro (same subjects, new questions) looks more like memorization than understanding.
Abnormally high score for its size
A small 7B model that beats 70B models on one specific benchmark and only that one: be wary—it is often targeted training (“benchmaxxing”) on that test set.
Dated benchmarks
“Live” evaluations (LiveCodeBench, time-stamped questions) work around the problem: only material after the model's cutoff date is scored.
Suspected memory leaks
Some tests inject canary variants (“canary strings”) or paraphrases to detect memorization. A large gap between the original and paraphrased question reveals contamination.

#Pitfall No. 2: saturation

A benchmark is saturated when the best models reach scores so high that it can no longer distinguish them. When everyone scores 96-99%, the remaining 3 points are measurement noise (ambiguous questions, labeling errors) rather than a real difference in capability. HumanEval (code) and MMLU (knowledge) are the canonical examples: they were very useful, but are no longer useful for separating the top of the rankings.

That is why tougher benchmarks keep appearing: MMLU-Pro replaces MMLU, GPQA Diamond and HLE take over for reasoning, and SWE-bench Verified replaces HumanEval for code. A benchmark has a useful lifespan; once the average level passes a certain point, remove it from your decision criteria.

!
Do not compare in the saturation zone
Choosing between two models at 97.1% and 97.8% on a saturated benchmark means choosing based on noise. Go down a level: look at a harder benchmark, or better yet, test them on your own tasks. The difference that matters to you is almost never in those 0.7 points.

#How to read a leaderboard without being misled

Here is the method to apply to any score table, whether it comes from a model announcement or a public ranking.

  1. 01
    Identify who publishes it
    A chart in a model announcement is marketing: it selects favorable benchmarks and an advantageous protocol (cherry-picking). A neutral, third-party leaderboard (Open LLM Leaderboard, LMArena, the official SWE-bench page) is far more reliable than a launch slide.
  2. 02
    Check the protocol
    0-shot or few-shot? With chain-of-thought? Which harness for agentic systems? Two scores can only be compared if the method is identical. A “self-reported” asterisk is worth less than a score reproduced by a third party.
  3. 03
    Look at the confidence intervals
    On LMArena, an Elo gap of less than ~15 points is noise. On low-sample benchmarks (GPQA Diamond, 198 questions), a handful of correct answers can move the score by several points. A ranking without a margin of error should be taken with a grain of salt.
  4. 04
    Cross-reference multiple benchmarks
    Never decide based on a single number. A solid model performs well across a varied family of tests, not just one peak. An isolated, abnormal peak is more suggestive of targeted optimization.
  5. 05
    Weight it according to YOUR use case
    Do you write code? Look at SWE-bench and LiveCodeBench, not MMLU. French? None of these benchmarks are in French—look for French evaluations or test them yourself. Conversation? LMArena is the priority. The best model “on average” is not necessarily the best one for you.

#Evaluate a local model on your own tasks

The logical conclusion from everything above: the most reliable benchmark for you is your own. It cannot be contaminated (your questions are not on the web), it is never saturated (you calibrate it to your difficult cases), and it measures exactly what you need. No heavy infrastructure is required: around twenty representative examples are enough to distinguish between two models.

  1. 01
    Build a mini test set
    Gather 15 to 30 prompts from your real use cases (emails to draft, extractions, questions about your documents, code snippets). For each one, note what a good answer must contain. Keep this set private and stable so you can compare models over time.
  2. 02
    Install the models to compare
    With Ollama, pull the candidate models. Ollama exposes an OpenAI-compatible API on http://localhost:11434, allowing calls to be automated.
  3. 03
    Automate calls
    A small script loops through your prompts and records each model's responses side by side in a file, so you can compare them later without being influenced by the model name.
  4. 04
    Record results blindly
    Review the answers without knowing which model produced what (shuffle the order). Score them using your own criteria: accuracy, adherence to the instructions, quality of the French, and absence of fabrication. This is your personal “Arena.”
  5. 05
    Also measure the real cost
    Quality is not everything: track speed (tokens per second) and VRAM usage. A 14B model in Q4_K_M (~9 GB) that fits on your RTX 4070 and responds quickly may outperform a 70B model (~40 GB) that is theoretically “better” but unusable on your system.
Terminal — preparing the models to compare
# Récupérer deux candidats via Ollama
ollama pull qwen3:14b
ollama pull gemma3:12b

# Vérifier qu'ils répondent (API locale sur le port 11434)
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3:14b",
  "prompt": "Résume ce texte en 3 puces : ...",
  "stream": false
}'
eval_local.py — compare two models on your prompts
import json, requests

OLLAMA = "http://localhost:11434/api/generate"
MODELES = ["qwen3:14b", "gemma3:12b"]

# Vos prompts réels — la clé d'une éval qui vous ressemble
prompts = [
    "Rédige un mail de relance poli à un client en retard de paiement.",
    "Extrais les dates et montants de ce texte : ...",
    "Explique la différence entre Q4_K_M et Q8_0 en 2 phrases.",
]

def interroger(modele, prompt):
    r = requests.post(OLLAMA, json={
        "model": modele, "prompt": prompt, "stream": False
    }, timeout=120)
    return r.json()["response"].strip()

resultats = []
for p in prompts:
    ligne = {"prompt": p}
    for m in MODELES:
        ligne[m] = interroger(m, p)
    resultats.append(ligne)

# À relire en aveugle, sans regarder la colonne du modèle
with open("comparaison.json", "w", encoding="utf-8") as f:
    json.dump(resultats, f, ensure_ascii=False, indent=2)
print("OK — comparaison.json généré, notez les réponses à froid.")
→
Tools if you want to industrialize
To go beyond a homemade script, tools such as lm-evaluation-harness (the academic standard, which powers the Open LLM Leaderboard) or promptfoo (focused on comparing prompts and models) can automate scoring. But for a personal choice, 20 manually scored examples beat any public score.

#Go further

These guides build on the concepts covered here, focusing on specialized benchmarks and the practical choice of a local model:

Code benchmarks in detail
“HumanEval Is Dead: Understanding LLM Code Benchmarks in 2026” explores SWE-bench, LiveCodeBench, and how to read code scores—the direct development-focused continuation of this guide.
Choosing your quantization
“GGUF Quantization in 2026: Q4_K_M vs Q5_K_M vs Q6_K” shows how to measure the real impact of quantization on quality—an in-house evaluation applied to a concrete case.
A complete model test
“Qwen 3 locally: complete test and real-world benchmarks” illustrates the local evaluation method (tokens/sec, French quality, VRAM) on a specific model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.