Local LLM benchmark: MMLU, HumanEval, AIME 2026
Un local LLM benchmark read outside its methodology, amounts to comparing two marathon times without the course elevation profile. Two models reported at 88 points on MMLU may differ by 16 points on MMLU-Pro and by 30 points on AIME 2026 depending on whether chain-of-thought training was used, the quantization format, the decoding protocol (cons@64 vs strict pass@1), and the exact version of the loaded dataset. This page is not a buying guide: it is an audit framework for the three historical suites (MMLU/MMLU-Pro, HumanEval/HumanEval+, AIME 2024/2025/2026), with published scores, reproducibility conditions, measured Q4 VRAM cost, tokens-per-second throughput on real configurations, and an inventory of biases that vendor press releases leave unsaid. The goal is practical: enable you to rerun a figure yourself with a fixed seed, not to arbitrate between proprietary models and open weights.
Why the 2026 evaluation framework invalidates pre-2024 benchmarks
The ecosystem shifted in two years. MMLU with 4 choices was already saturated from Llama 3.1 405B at 88.6, and expert-activation architectures (DeepSeek V3, Qwen 3, GLM-5.1) are now nearing the ceiling; reading an MMLU score to the nearest half point has become informationally meaningless. Three structural shifts have reshuffled the matrix:
- Reasoning models trained with RL (DeepSeek R1, Ring-1T, gpt-oss 120B with activatable thinking) that turn AIME and MATH-500 into leading discriminators. On AIME 2024, the gap Llama 3.1 405B (~23 pass@1) vs DeepSeek R1 671B (~79 pass@1) reaches 55 points — a delta impossible to observe on MMLU.
- Total parameters vs. active parameters breakdown. Qwen 3 235B-A22B activates only 22B parameters per token, MiniMax-M2.7 follows the same MoE logic. The benchmark should be interpreted at equivalent inference cost (tokens/second × VRAM), not by the displayed total parameter count.
- Anti-contamination suites. LiveBench refreshes its items every month, LiveCodeBench timestamps Codeforces problems to isolate those created after the pre-training cutoff, and SimpleBench measures common-sense robustness where MMLU plateaus. The HuggingFace Open LLM Leaderboard v2 explicitly migrated to MMLU-Pro, GPQA, MUSR, and IFEval for this reason.
Reading an isolated score no longer tells you anything: only the matrix {MMLU-Pro, HumanEval+, AIME 2026, SWE-bench Verified, RULER 128k, GPQA Diamond} provides signal. The guide /guide/matrice-benchmarks-2026 explains the recommended weighting.
Methodology: what each benchmark suite actually measures
Le MMLU LLM (Massive Multitask Language Understanding, Hendrycks et al. 2021) covers 57 disciplines (medicine, law, ethics, elementary mathematics) in 4-choice multiple-choice questions, using the standard 5-shot protocol. Reference repo: GitHub Hendrycks, paper arXiv:2009.03300. Beyond 85%, it was superseded by MMLU-Pro (10 choices, multi-step reasoning, 12,032 items) on HuggingFace TIGER-Lab, with a typical drop of 15 to 25 points compared with standard MMLU. Emerging variant: MMLU-Redux (arXiv paper:2406.04127) corrects about 6.5% of incorrectly recognized or ambiguous items in the original set—an often-overlooked point that explains why two identical runs can differ by 1 to 2 points depending on the loaded version. For the filters and evaluation procedure, see /guide/mmlu-pro-protocole-strict.
HumanEval, published by OpenAI (arXiv:2107.03374), evaluates 164 Python problems with function signatures and unit tests. The standard metric is pass@1 mais pass@10 et pass@100 require 10 and 100 samples per problem, respectively, multiplying the evaluation cost accordingly. Its successor HumanEval+ (evalplus) adds about 80× more mutation-generated test cases; overfit models typically lose 5 to 10 points between the two versions, directly revealing the degree of overfitting to the original tests. To go further with real-world code, SWE-bench Verified measures the resolution of authentic GitHub issues (500 human-validated tickets) and LiveCodeBench maintains a stream of timestamped Codeforces problems that can be filtered by date. The complete extraction procedure for HumanEval+ is covered in /guide/humaneval-plus-protocole.
AIME 2024 LLM corresponds to the 15 annual problems of the American Invitational Mathematics Examination (two sessions, AIME I and II), with integer answers between 0 and 999. Having become central with reasoning models, it strongly differentiates between architectures with or without chain-of-thought trained with RL. The 2024 set is referenced in the model card DeepSeek R1. AIME 2025 (March 2025), followed by AIME 2026 (February 2026), now serves as a fresh evaluation because the 2024 problems have been extensively indexed. Technical reports from late 2025 often use cons@64 (majority voting over 64 samples, temperature 0.6, top-p 0.95), which can inflate the raw score by 10 to 15 points compared with a strict pass@1 at temperature 0.
Beyond these three pillars, the 2026 grid adds GPQA Diamond (198 doctoral-level questions validated by experts), MATH-500 (Olympiad problems), MuSR (multi-step narrative reasoning), RULER (long-context retrieval up to 1M tokens), IFEval (strict instruction following), BFCL v3 (Berkeley Function Calling Leaderboard) for structured tool calling. Serious technical reports now publish this complete matrix rather than an isolated MMLU score.
Published scores: comparative evaluation on equivalent hardware
The figures below come from official technical specifications or Hugging Face reports. They still need to be reproduced via EleutherAI's lm-evaluation-harness, the only authoritative independent protocol in the open-source community.
Tier S — frontier reasoning models (> 600B parameters) :
- DeepSeek R1 671B : MMLU ~90.8, HumanEval ~96, AIME 2024 ~79.8 pass@1. MoE architecture with 37B active parameters, MIT license, reasoning mode by default.
- DeepSeek V3.2 : MMLU ~88.5, estimated HumanEval ~93, 128k window, general-purpose version without long thinking—much lower inference cost than R1.
- DeepSeek V4 Pro 1.6T : 1600B parameters, 1M-token context window, very high long-context RULER scores, realistically deployable only on an 8×H200 cluster.
- Kimi K2.6 : estimated MMLU ~89, AIME 2024 ~74 according to Moonshot. Successor to Kimi K2.5, same ~600 GB Q4 VRAM, Modified MIT license.
- GLM-5.1 : 744B MoE parameters, competitive MMLU, AIME 2024 scores announced by Z.AI and requiring independent confirmation.
- Mistral Large 3 675B : Apache 2.0, MMLU at the DeepSeek V3 level, particularly strong in French and with native function calling.
- Ring-1T et Ling 2.6 1T from Ant Group / inclusionAI: 1000B parameters, explicit-reasoning focus.
Tier A — runnable on 1 to 4 high-end GPUs (100B to 400B) :
- Llama 3.1 405B Instruct : MMLU 88.6, HumanEval 89.0 (Meta report), historical dense baseline.
- Llama 4 Maverick 400B : MoE with a 1M-token window, Q4 deployment ~240 GB.
- Qwen 3 235B-A22B : MMLU-Pro ~75 estimated, HumanEval ~92. Only 22B active parameters: excellent quality-to-VRAM ratio on 2×H100.
- DeepSeek V4 Flash 284B : 1M-token long-context variant, retrieval-heavy target.
- gpt-oss 120B : OpenAI release under Apache 2.0, high AIME 2024 performance for its size thanks to reasoning mode, deployable on 1×H100 or a 192 GB M3 Ultra Mac Studio.
- Nemotron 3 Super 120B : NVIDIA, optimized for TensorRT-LLM, competitive on MMLU.
- Mistral Small 4 119B : 256k tokens, Apache 2.0, open alternative to Mistral Medium 3.5.
- ERNIE 4.5 300B-A47B : Baidu, Apache 2.0, good results in Chinese and a 131k context window.
- Hunyuan Large 2.0 : Tencent, MoE 389B/52B active, 262k context window.
Tier B — single workstation (40 to 80 GB VRAM) :
- Llama 3.3 70B Instruct : MMLU ~86.0, HumanEval ~88.4 (Meta model card).
- DeepSeek R1 Distill Llama 70B : reasoning distillate, AIME 2024 ~70 pass@1 — record reasoning/VRAM ratio for 40 GB.
- Qwen3-Coder-Next 80B-A3B : code-specialized, 3B active MoE, very high throughput in SGLang.
- Llama 3.1 Nemotron 70B : RLHF tuning NVIDIA, high Arena scores.
- Tülu 3 70B : fully open post-training recipe from Allen AI (github.com/allenai/open-instruct).
- Hunyuan-A13B Instruct : 80B total, 13B active, an interesting MoE alternative.
Close comparisons available at /compare/deepseek-r1-distill-llama-70b-vs-llama33-70b et /compare/qwen3-235b-a22b-vs-llama-3-1-405b.
Detailed hardware costs: VRAM, quantization, measured throughput
A raw score means nothing without the associated inference cost. The values below correspond to the Q4_K_M format of llama.cpp. Converting to Q5_K_M, Q8_0, or FP16 multiplies VRAM by approximately 1.25, 2, and 4. The calculator quelllm.fr/configurateur GPU budget filter.
| Model | Q4 VRAM | Target configuration | Indicative Q4 throughput |
|---|---|---|---|
| DeepSeek V4 Pro 1.6T | ~960 GB | 8×H200 141 GB, TPU v5p pod | 15-20 tok/s tensor-parallel |
| MiMo V2.5 Pro 1020B | ~595 GB | 8×H100 80 GB tensor-parallel | 20–30 tok/s |
| Mistral Large 3 675B | ~405 GB | 6×H100 80 GB or 4×H200 | 25-35 tok/s |
| Llama 3.1 405B | ~240 GB | 4×A100 80 GB or 3×H100 | 20–30 tok/s |
| Qwen 3 235B-A22B | ~142 GB | 2×H100 or 1×H200 141 GB | 40–60 tok/s (MoE) |
| Hunyuan Large 2.0 | ~245 GB | 3×H100 80 GB | 25–30 tok/s |
| gpt-oss 120B | ~70 GB | 1×H100, Mac Studio M3 Ultra 192 GB | 35–50 tok/s |
| Mistral Small 4 119B | ~72 GB | 1×H100, A100 80 GB | 30–45 tok/s |
| Llama 3.3 70B | ~40 GB | RTX A6000 48 GB, 2×RTX 4090 | 15–25 tok/s on 2×4090 |
| DBRX Instruct 132B | ~76 GB | 1×H100, 2×A100 40 GB | 30-40 tok/s |
The figures are rough orders of magnitude for single-batch operation; multi-tenant deployments with vLLM or SGLang achieve up to 5× more aggregate throughput through continuous batching and prefix caching. On Mac Studio M3 Ultra (800 GB/s memory bandwidth), Llama 3.3 70B in Q4_K_M reaches 8 to 12 tok/s in single-stream depending on the effective context length. MoE architectures (Mixtral 8x22B, Qwen 3, DeepSeek V3) benefit particularly from SGLang thanks to batched shared expert routing. The guide /guide/quantization-q4-q5-q8 details the quality/VRAM trade-offs and /guide/inference-vllm-vs-llamacpp compares engines in depth. To benchmark your own setup, see /guide/bench-tokens-par-seconde.
Three biases that distort ranking interpretation
- Dataset contamination. MMLU and HumanEval have circulated since 2021; some models saw the questions during pretraining, which artificially inflates the scores. The work LiveCodeBench provides time-based tracking of HumanEval issues to isolate those that appeared after the pretraining cutoff. Detection of contamination through n-gram overlap (method arXiv:2311.04850) shows that up to 12% of MMLU items appear verbatim in indexed web corpora.
- Decoding variance. An AIME pass@1 score calculated with temperature 0.6 and 64 samples (cons@64, majority voting) may differ by 10 points from a strict pass@1 at temperature 0. On GSM8K, the self-consistency gap reaches 5 to 8 points depending on the number of samples. Always verify
temperature,top_p,max_tokens, number of samples, and aggregation rule in the model card or technical report. - Specific prompts. The DeepSeek, Qwen, and Mistral reports often use optimized system prompts, sometimes with manually curated few-shot examples. Reproducing the figures with raw zero-shot typically yields 3 to 8 points less, sometimes more on reasoning suites. For AIME, adding a “Let's think step by step” prefix can shift the score by 4 to 6 points on models without trained-in thinking.
User-task-oriented benchmarks — LMArena (formerly Chatbot Arena), LiveBench, SWE-bench Verified — provide a less manipulable complement. SWE-bench remains the benchmark for evaluating an agentic model on real-world code: DeepSeek V3.2 et Qwen3-Coder-Next 80B-A3B achieve scores above 40%, with values to be confirmed on the updated leaderboard. BFCL v3 adds the structured tool-calling capability useful for agentic deployments.
Reproduce a benchmark end to end: standard procedure
To produce a defensible AIME 2026 score on, for example, DeepSeek R1 Distill Llama 70B :
- Retrieve the weights on HuggingFace (
# HuggingFace : deepseek-ai/DeepSeek-R1-Distill-Llama-70B) then convert to GGUF Q4_K_M viallama.cpp/convert_hf_to_gguf.pyif you’re targeting CPU/Apple, or stay with safetensors for vLLM. - Set the hyperparameters : temperature 0.6, top-p 0.95, max_tokens 32768 (the model needs to reason at length), seed recorded for reproducibility.
- Load the 2026 AIME set from an updated Hugging Face dataset. Check that the publisher did not accidentally include the solution in the prompt—a common issue with recent forks.
- Generate 64 completions per problem for cons@64, and 1 completion at temperature 0 for strict pass@1.
- Extract the final answer via regex on the tag
\boxed{}or the last sequence of 1 to 3 digits. Handle cases where the model emits multiple tags<thinking>nested. - Compare with official solutions AIME (integer 0–999), publish script and seed.
Allow approximately 6 to 10 hours on 2×H100 for a 70B model at cons@64 on the 15 AIME 2026 problems. The guide /guide/reproduire-benchmark-aime details the pitfalls of answer extraction, particularly handling tags <thinking> among reasoning models and normalization of output fractions.
Quantified case studies: three concrete scenarios
Scenario A — university lab with 4×A100 80 GB for evaluating mathematical reasoning. Goal: produce a reproducible AIME 2024/2025/2026 curve. Good choice: DeepSeek R1 Distill Llama 70B in Q8_0 (140 GB, two GPUs), compared with Llama 3.3 70B without reasoning on the two remaining GPUs to measure the pure contribution of chain-of-thought. Cost of one complete cons@64 run on the 45 problems (AIME 2024 + 2025 + 2026): approximately 30 GPU hours. See /compare/deepseek-r1-distill-llama-70b-vs-llama33-70b.
Scenario B — startup with 1×H100 80 GB for benchmarking a coding assistant. Three candidates head-to-head: Qwen3-Coder-Next 80B-A3B, gpt-oss 120B et Mistral Small 4 119B. Protocol: HumanEval+ pass@1, MBPP+ pass@1, LiveCodeBench filtered to problems after October 2024, SWE-bench Verified Lite (50 issues). The MoE throughput of Qwen3-Coder-Next (3B active) typically enables 3× more candidates per hour than dense models with comparable VRAM, reducing the value of aggressive quantization.
Scenario C — data team with a Mac Studio M3 Ultra 192 GB. The 800 GB/s memory bandwidth remains the limiting factor. Llama 3.3 70B Q5_K_M fits comfortably and delivers 6 to 10 tok/s. Qwen 3 235B-A22B in Q4 just fits within 142 GB, at 4–6 tok/s but with superior quality. gpt-oss 120B in Q4 occupies 70 GB and leaves room for the long-context KV cache. See /guide/llm-mac-studio-m3-ultra.
Scenario D — benchmark a regression between two fine-tuned checkpoints. Typical industrial use case: a LoRA fine-tune on 10k business-domain examples on Llama 3.3 70B. The risk is catastrophic regression in general capabilities. Minimal protocol: MMLU (5-shot), IFEval, GSM8K, HumanEval+ before and after fine-tuning, with the same seed. A drop of more than 2 points on 2 of the 4 suites signals overfitting. See /guide/fine-tuning-sans-regression.
Exact licenses and commercial-use terms
Le local LLM benchmark is useless if the license prohibits the intended deployment. Four families dominate the 2026 open-weights ecosystem:
- MIT / Apache 2.0 (free commercial use, attribution required): DeepSeek R1 671B, DeepSeek V3.2, Mistral Large 3, Mixtral 8x22B, Qwen 3 235B-A22B, gpt-oss 120B, Ring-1T, Ling 2.6 1T, GLM-5.1, Rakuten AI 3.0, Snowflake Arctic Instruct, MiniMax-M2.7, Step 3.5 Flash. Comfortable range for B2B SaaS without a user threshold.
- Llama Community License (restrictions beyond 700M monthly MAUs): Llama 3.1 405B, Llama 4 Maverick 400B, Llama 4 Scout 109B, Llama 3.3 70B, Llama 3.1 70B. Read before any mass-market product with a large audience.
- Specific vendor licenses : Tencent Hunyuan License for Hunyuan Large 2.0 et Hunyuan-A13B, Qwen License for Qwen 2.5 72B, Pangu Model License for Pangu Pro MoE 72B, NVIDIA Open Model License for Nemotron 3 Super 120B, Databricks Open Model License for DBRX.
- Noncommercial : CC-BY-NC 4.0 for Command R+ 104B. For academic research only; not suitable for paid production.
For a detailed license-by-license comparison, see /guide/licences-llm-open-source.
Choose a model based on your target benchmark profile
- Mathematical reasoning (AIME, MATH-500, GPQA Diamond) : DeepSeek R1 671B in Q4, or its more accessible distillate DeepSeek R1 Distill Llama 70B. See /meilleur-llm/raisonnement et /compare/deepseek-r1-671b-vs-qwen3-235b-a22b.
- Code generation (HumanEval+, MBPP+, SWE-bench Verified) : Qwen3-Coder-Next 80B-A3B for a VRAM/quality tradeoff, Mistral Large 3 675B for the high end. See /meilleur-llm/code et /compare/qwen3-coder-next-vs-mistral-large-3.
- Long context (RULER 128k-1M, needle-in-a-haystack) : MiMo V2.5, DeepSeek V4 Flash 284B, Llama 4 Maverick 400B. Llama 4 Scout 109B claims 10M tokens, to be confirmed on real-world retrieval.
- Multilingual with French : Mistral Large 3 675B, Mistral Medium 3.5 128B, Qwen 3 235B-A22B. For Arabic and MENA, Jais Adapted 70B Chat. For Japanese, Rakuten AI 3.0. See /meilleur-llm/francais.
- Vision (MMMU, MathVista, DocVQA) : Qwen 3 VL 235B-A22B, Qwen 2.5 VL 72B, LLaVA-OneVision 72B, Molmo 72B. See /meilleur-llm/vision.
- Limited VRAM budget (< 48 GB) : Llama 3.3 70B, Llama 3.1 Nemotron 70B, Tülu 3 70B in Q4. See /meilleur-llm/vram-48go.
- Tool calling (BFCL v3, Tau-Bench) : Mistral Large 3, Qwen 3 235B-A22B, Hunyuan-A13B. See /meilleur-llm/agents.
FAQ
Q: What is the difference between MMLU and MMLU-Pro?
MMLU has 4 choices per question and tops out at around 88-90% for the best open-weight models, making it poorly discriminative. MMLU-Pro increases this to 10 choices, eliminates ambiguous questions, and adds items requiring multi-step reasoning. Scores typically drop by 15 to 25 points, restoring its discriminative power. To compare modern models such as DeepSeek V3.2 or Qwen 3 235B-A22B, MMLU-Pro is now more relevant than standard MMLU.
Q: How can you reproduce a HumanEval score locally?
Install evalplus, load the model with vLLM or llama.cpp, generate 1 to 200 completions per problem depending on the target metric, then run the Python tests. Budget one GPU-day for a 70B model at pass@1 across all 164 problems, and more for pass@100. A score at temperature 0 differs from a score averaged across multiple samples: always document the hyperparameters and publish the script to enable cross-comparison.
Q: Should you prefer Q4, Q5, or Q8 for a reliable benchmark?
Q4_K_M typically loses 1 to 3 points on MMLU compared with FP16, sometimes more on AIME, where long-form reasoning is sensitive to cumulative errors. Q5_K_M and Q6_K reduce the gap to less than one point. Q8_0 is nearly indistinguishable from FP16 on most suites. For honest benchmarking, always specify the quantization used. The guide /guide/quantization-q4-q5-q8 details the measured drops for each format.
Q: Are the AIME 2024 scores contaminated?
The AIME 2024 questions were published online in February 2024 and indexed by crawlers soon afterward. Models with a data cutoff after mid-2024 may therefore have seen the solutions. For this reason, AIME 2025 and AIME 2026 are prioritized in recent evaluations, provided the models were not exposed to them. Always check the pre-training cutoff date announced by the publisher in the model card and cross-check it against LiveBench to confirm.
Q: Which model for a system with 24 GB of VRAM?
24 GB rules out 70B+ models in Q4, which require at least 40 GB. Viable options are 30B to 40B models in Q4, or 70B models in aggressive quantizations such as IQ2_XXS, at the cost of a significant quality loss (5 to 10 points on MMLU). To stay within an acceptable quality range, target 14B to 32B models. See the configurator at /configurateur for VRAM filtering.
Q: Why do scores differ between HuggingFace and vendor reports?
Editors optimize prompts, use specialized decoding (cons@64, majority voting), and sometimes select the best runs across multiple seeds. Independent leaderboards such as HuggingFace Open LLM Leaderboard require a uniform protocol (fixed zero-shot or few-shot, standardized prompts). A gap of 3 to 10 points between the two sources is expected. In production, the reproducible protocol matters, not the maximum score.
Conclusion
Un local LLM benchmark read without the context of the license, quantization, and decoding protocol, is misleading. The robust approach is to cross-check at least MMLU-Pro, HumanEval+, AIME 2025/2026, and an agentic benchmark such as SWE-bench Verified, verifying the target VRAM and license before making any commitment. The scope here is intentionally limited to measurement: what matters on this page is the ability to reproduce a published figure, understand its margin of error, and validate that it runs on your hardware. To filter indexed models by your GPU budget and use case, launch the BestLLMfor configurator or browse the full catalog.