MMLU-Pro: the knowledge benchmark for local LLMs

Le MMLU-Pro benchmark has become the benchmark for evaluating the depth of knowledge and reasoning ability of open-weights models run locally. Published in 2024 by TIGER-Lab, this MMLU-Pro benchmark fixes the weaknesses of the original MMLU (saturation, ambiguous questions, overly guessable multiple-choice options) by introducing 12,032 questions with 10 options across 14 disciplines. This article details the protocol, observed scores on models in the quelllm.fr catalog, the impact of quantization on results, and how to reproduce the benchmark at home on a consumer GPU or workstation.

Why MMLU-Pro replaced MMLU in serious evaluations

The classic MMLU (Hendrycks et al., 2020) tops out at around 88–90% for the best models, making it nearly impossible to distinguish between frontier models. The arXiv paper 2406.01574 introduces three major fixes:

Result: the average score drops by 15 to 25 points compared with MMLU. A model that scores 86% on MMLU typically drops to around 66% on MMLU-Pro. This downward compression makes the metric meaningful again for comparing modern open-weight models.

The complete dataset is hosted on HuggingFace TIGER-Lab/MMLU-Pro and the reference evaluation is performed via the official GitHub repository.

MMLU-Pro scores observed in the quelllm.fr catalog

The scores below come from the official TIGER-Lab leaderboard and technical reports from publishers. Values marked “estimated” come from partial community measurements that require confirmation.

Differences between categories (law, medicine, engineering) can reach 20 points for the same model: an average score often masks a domain-specific weakness.

Impact of quantization on the LLM MMLU score

Quantization reduces VRAM usage but lowers scores. Empirical measurements on the models in the catalog give these approximate ranges (sources: llama.cpp issues, Unsloth reports):

Pour Llama 3.3 70B Instruct, for example, we see a drop from ~68% in FP16 to ~64% in Q4_K_M on the "law" subset. On Qwen 2.5 72B Instruct, the gap is smaller (1.5 points). The rule of thumb: Q5_K_M remains the best compromise for serious evaluations, while Q4 is suitable for everyday use.

See the GGUF quantization guide for format details.

VRAM and hardware costs for running the knowledge benchmark

Reproducing MMLU-Pro locally requires loading the model into memory and generating approximately 12,032 responses (with CoT, this amounts to 5 to 10 million output tokens). Hardware cost by tier:

A full benchmark on a 70B model in Q4 takes 4 to 8 hours on RTX 6000 Ada, depending on throughput (~25-35 tokens/sec). To calibrate your setup, the tool /configurateur recommends hardware suited to a target VRAM budget.

How to reproduce the benchmark locally

The standard protocol uses vllm or llama.cpp on the backend, using the official Python script from the TIGER-AI-Lab repository.

Main steps:

  1. Dataset download depuis HuggingFace (about 12 MB).
  2. Model configuration in chat completion mode with temperature 0 and top-p 1. The system prompt must include 5 few-shot examples per category (CoT enabled).
  3. Start the evaluation by discipline (14 categories: biology, business, chemistry, computer_science, economics, engineering, health, history, law, math, other, philosophy, physics, psychology).
  4. Response parsing : extraction of the final letter (A to J) via regex on the last line.
  5. Score calculation weighted by category or overall average.

Full details in the GitHub README. To compare two models side by side, see /compare/llama33-70b-vs-qwen25-72b or view the overall ranking at /meilleur-llm/raisonnement.

Limitations of MMLU-Pro and Additional Benchmarks

No benchmark is sufficient on its own. MMLU-Pro measures academic knowledge and structured reasoning, but does not cover:

A robust test suite combines at least MMLU-Pro + HumanEval + GSM8K + an agentic benchmark.

FAQ

Q: What is the practical difference between MMLU and MMLU-Pro?

MMLU has 4 options per question and saturates above 88%. MMLU-Pro increases this to 10 options, filters ambiguous questions, and incorporates multi-step reasoning. Result: scores drop by 15 to 25 points, restoring discrimination. For 2025+ use, MMLU-Pro is now the benchmark in vendors' technical reports.

Q: Should I enable chain-of-thought to evaluate a model?

Yes for comparative benchmarks, because the official MMLU-Pro protocol uses it. Without CoT, scores drop by 5 to 15 points depending on the model. On DeepSeek R1 671B, reasoning mode gains ~8 points over a direct call. To measure production relevance without CoT, run the evaluation separately in "answer only" mode.

Q: Can I run the knowledge benchmark on a Mac M4?

Yes, via llama.cpp or MLX. A 128 GB Mac Studio M4 Max runs Llama 3.3 70B Instruct in Q4 (~40 GB) at 8-12 tokens/sec, or 6 to 10 hours for the complete benchmark. Models above 100B require an M4 Ultra 192 GB or very slow SSD offloading. See the Mac Apple Silicon guide for the recommended configurations.

Q: Is the MMLU-Pro score reliable for choosing a business model?

Partially. MMLU-Pro reflects general-purpose capability well, but an overall score of 70% can hide 55% in law and 82% in physics. If your use case targets a specific discipline, look at the category score published on the leaderboard. For French-language coding or legal use, supplement it with HumanEval-FR and an internal test on your own data.

Q: What license applies to the top models in the benchmark?

The top open-weight scores come from models under permissive licenses: DeepSeek V3 671B (DeepSeek License), DeepSeek R1 671B (MIT), Qwen 3 235B-A22B (Apache 2.0), Mistral Large 3 675B (Apache 2.0). Note: Llama 3.1 405B Instruct is under Llama 3.1 Community (restrictions above 700M MAU). Check each license before a commercial deployment.

Q: How much does a complete evaluation on a rented GPU cost?

At €2/hour for an H100 80 GB via bare-metal cloud, a complete MMLU-Pro benchmark on a 70B in Q4 costs €8 to €16 depending on throughput. For a 405B in Q4 requiring 4 H100s, expect €80 to €160. This is significantly cheaper than buying the hardware, but local inference remains relevant for repeated evaluations or sensitive data.

Conclusion

Le MMLU-Pro benchmark offers today's most discriminating measure of knowledge and reasoning for open-weight LLMs. Combined with specialized benchmarks (HumanEval, AIME, LongBench), it provides an objective framework for choosing a model. To identify the hardware configuration that lets you run the top-ranked models, use the quelllm.fr configurator or explore the 249 models indexed on the full catalog.

Article published on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.