MMLU-Pro: the knowledge benchmark for local LLMs
Le MMLU-Pro benchmark has become the benchmark for evaluating the depth of knowledge and reasoning ability of open-weights models run locally. Published in 2024 by TIGER-Lab, this MMLU-Pro benchmark fixes the weaknesses of the original MMLU (saturation, ambiguous questions, overly guessable multiple-choice options) by introducing 12,032 questions with 10 options across 14 disciplines. This article details the protocol, observed scores on models in the quelllm.fr catalog, the impact of quantization on results, and how to reproduce the benchmark at home on a consumer GPU or workstation.
Why MMLU-Pro replaced MMLU in serious evaluations
The classic MMLU (Hendrycks et al., 2020) tops out at around 88–90% for the best models, making it nearly impossible to distinguish between frontier models. The arXiv paper 2406.01574 introduces three major fixes:
- 10 options per question instead of 4 : the probability of random success drops from 25% to 10%, broadening the range of scores.
- Manual filtering of ambiguous questions : approximately 5,000 questions from the original MMLU were removed or reworded by experts.
- Adding questions with a high reasoning coefficient : integration of STEM problems (TheoremQA, SciBench) that require multi-step reasoning rather than simple memory retrieval.
Result: the average score drops by 15 to 25 points compared with MMLU. A model that scores 86% on MMLU typically drops to around 66% on MMLU-Pro. This downward compression makes the metric meaningful again for comparing modern open-weight models.
The complete dataset is hosted on HuggingFace TIGER-Lab/MMLU-Pro and the reference evaluation is performed via the official GitHub repository.
MMLU-Pro scores observed in the quelllm.fr catalog
The scores below come from the official TIGER-Lab leaderboard and technical reports from publishers. Values marked “estimated” come from partial community measurements that require confirmation.
- DeepSeek V3 671B : 75.9% — the best general-purpose open-weights score recorded as of the end of 2024. Full profile at /modele/deepseek-v3-671b.
- DeepSeek R1 671B : 84.0% (reasoning enabled)—~8-point gain thanks to internal chain-of-thought. Details on /modele/deepseek-r1-671b.
- Qwen 3 235B-A22B : 72.1% estimated — MoE architecture, see /modele/qwen3-235b-a22b.
- Llama 3.1 405B Instruct : 73.3% — Meta reference, spec sheet /modele/llama-3-1-405b.
- Llama 3.3 70B Instruct : 68.9% — best performance/VRAM ratio for 40 GB workstations, spec sheet /modele/llama33-70b.
- Mistral Large 3 675B : estimated at 71.5% — see /modele/mistral-large-3.
- gpt-oss 120B : estimated at 70.2%—distilled OpenAI variant, page /modele/gpt-oss-120b.
- Qwen 2.5 72B Instruct : 65.4% — solid multilingual baseline, spec sheet /modele/qwen25-72b.
- Mixtral 8x22B Instruct : 56.2% — historical but still relevant for comparison, spec sheet /modele/mixtral-8x22b.
Differences between categories (law, medicine, engineering) can reach 20 points for the same model: an average score often masks a domain-specific weakness.
Impact of quantization on the LLM MMLU score
Quantization reduces VRAM usage but lowers scores. Empirical measurements on the models in the catalog give these approximate ranges (sources: llama.cpp issues, Unsloth reports):
- FP16 → Q8_0 : loss of < 0.5 points on MMLU-Pro, generally not significant.
- Q8 → Q5_K_M : a loss of 1 to 2 points depending on the subject (mathematics is affected more).
- Q5 → Q4_K_M : a loss of 2 to 4 points, more pronounced on multi-step reasoning questions.
- Q4 → Q3 : possible collapse (> 5 points), to avoid for demanding use cases.
Pour Llama 3.3 70B Instruct, for example, we see a drop from ~68% in FP16 to ~64% in Q4_K_M on the "law" subset. On Qwen 2.5 72B Instruct, the gap is smaller (1.5 points). The rule of thumb: Q5_K_M remains the best compromise for serious evaluations, while Q4 is suitable for everyday use.
See the GGUF quantization guide for format details.
VRAM and hardware costs for running the knowledge benchmark
Reproducing MMLU-Pro locally requires loading the model into memory and generating approximately 12,032 responses (with CoT, this amounts to 5 to 10 million output tokens). Hardware cost by tier:
- 24 GB VRAM (RTX 4090) : limited to ~30B models in Q4; no access to large open-weight models.
- 48 GB VRAM (2× RTX 4090 or RTX 6000 Ada) : enables Llama 3.3 70B Instruct in Q4 (~40 GB), Qwen 2.5 72B Instruct (~42 GB), Qwen3-Coder-Next 80B-A3B (~48 GB).
- 96 GB VRAM (2× RTX 6000 Ada) : opens Mixtral 8x22B Instruct (~82 GB), gpt-oss 120B (~70 GB), Mistral Small 4 (~72 GB).
- 160-200 GB VRAM (4-5× RTX 6000 Ada or H100) : Llama 3.1 405B Instruct in Q4 (~240 GB with CPU offloading), Qwen 3 235B-A22B (~142 GB).
- 400+ GB VRAM (H100/H200 cluster) : DeepSeek V3 671B, DeepSeek R1 671B, Kimi K2.5 and beyond.
A full benchmark on a 70B model in Q4 takes 4 to 8 hours on RTX 6000 Ada, depending on throughput (~25-35 tokens/sec). To calibrate your setup, the tool /configurateur recommends hardware suited to a target VRAM budget.
How to reproduce the benchmark locally
The standard protocol uses vllm or llama.cpp on the backend, using the official Python script from the TIGER-AI-Lab repository.
Main steps:
- Dataset download depuis HuggingFace (about 12 MB).
- Model configuration in chat completion mode with temperature 0 and top-p 1. The system prompt must include 5 few-shot examples per category (CoT enabled).
- Start the evaluation by discipline (14 categories: biology, business, chemistry, computer_science, economics, engineering, health, history, law, math, other, philosophy, physics, psychology).
- Response parsing : extraction of the final letter (A to J) via regex on the last line.
- Score calculation weighted by category or overall average.
Full details in the GitHub README. To compare two models side by side, see /compare/llama33-70b-vs-qwen25-72b or view the overall ranking at /meilleur-llm/raisonnement.
Limitations of MMLU-Pro and Additional Benchmarks
No benchmark is sufficient on its own. MMLU-Pro measures academic knowledge and structured reasoning, but does not cover:
- The code : use HumanEval, MBPP, or LiveCodeBench. The model Qwen3-Coder-Next 80B-A3B is designed for this category (fiche).
- Advanced mathematics : AIME, MATH, and GSM8K remain essential. DeepSeek R1 671B shines especially here.
- The multilingual model : MGSM, Multilingual MMLU, Belebele for French. Mistral Large 3 675B et Rakuten AI 3.0 are better evaluated on these dimensions.
- The long context : RULER, LongBench. Llama 4 Scout 109B (10M context tokens) or MiMo V2.5 Pro (1M tokens) are the benchmarks here.
- L'agentique : SWE-bench, BFCL, AgentBench.
A robust test suite combines at least MMLU-Pro + HumanEval + GSM8K + an agentic benchmark.
FAQ
Q: What is the practical difference between MMLU and MMLU-Pro?
MMLU has 4 options per question and saturates above 88%. MMLU-Pro increases this to 10 options, filters ambiguous questions, and incorporates multi-step reasoning. Result: scores drop by 15 to 25 points, restoring discrimination. For 2025+ use, MMLU-Pro is now the benchmark in vendors' technical reports.
Q: Should I enable chain-of-thought to evaluate a model?
Yes for comparative benchmarks, because the official MMLU-Pro protocol uses it. Without CoT, scores drop by 5 to 15 points depending on the model. On DeepSeek R1 671B, reasoning mode gains ~8 points over a direct call. To measure production relevance without CoT, run the evaluation separately in "answer only" mode.
Q: Can I run the knowledge benchmark on a Mac M4?
Yes, via llama.cpp or MLX. A 128 GB Mac Studio M4 Max runs Llama 3.3 70B Instruct in Q4 (~40 GB) at 8-12 tokens/sec, or 6 to 10 hours for the complete benchmark. Models above 100B require an M4 Ultra 192 GB or very slow SSD offloading. See the Mac Apple Silicon guide for the recommended configurations.
Q: Is the MMLU-Pro score reliable for choosing a business model?
Partially. MMLU-Pro reflects general-purpose capability well, but an overall score of 70% can hide 55% in law and 82% in physics. If your use case targets a specific discipline, look at the category score published on the leaderboard. For French-language coding or legal use, supplement it with HumanEval-FR and an internal test on your own data.
Q: What license applies to the top models in the benchmark?
The top open-weight scores come from models under permissive licenses: DeepSeek V3 671B (DeepSeek License), DeepSeek R1 671B (MIT), Qwen 3 235B-A22B (Apache 2.0), Mistral Large 3 675B (Apache 2.0). Note: Llama 3.1 405B Instruct is under Llama 3.1 Community (restrictions above 700M MAU). Check each license before a commercial deployment.
Q: How much does a complete evaluation on a rented GPU cost?
At €2/hour for an H100 80 GB via bare-metal cloud, a complete MMLU-Pro benchmark on a 70B in Q4 costs €8 to €16 depending on throughput. For a 405B in Q4 requiring 4 H100s, expect €80 to €160. This is significantly cheaper than buying the hardware, but local inference remains relevant for repeated evaluations or sensitive data.
Conclusion
Le MMLU-Pro benchmark offers today's most discriminating measure of knowledge and reasoning for open-weight LLMs. Combined with specialized benchmarks (HumanEval, AIME, LongBench), it provides an objective framework for choosing a model. To identify the hardware configuration that lets you run the top-ranked models, use the quelllm.fr configurator or explore the 249 models indexed on the full catalog.