ARC-AGI 2: testing open-source LLM reasoning
Le ARC-AGI 2 benchmark became the benchmark in 2025 for measuring generalization and abstract reasoning in open-weight LLMs, far beyond classic multiple-choice tests such as MMLU. Designed by François Chollet and the Arc Prize Foundation, the ARC-AGI 2 benchmark subjects models to novel visual puzzles where memorization and brute force are no longer enough. This article reviews the evaluation protocol, the highest-ranked open-source models in the BestLLMfor catalog, their VRAM hardware costs, and best practices for reproducing these tests locally.
What is ARC-AGI 2 and why it changes the game
ARC-AGI 2 is the second iteration of the Abstraction and Reasoning Corpus, published as part of theArc Prize 2025 with a one-million-dollar prize. The benchmark contains approximately 1,000 public training tasks, 400 public evaluation tasks, and 100 private final-test tasks. Each task consists of a colored input grid and an output grid, accompanied by 2 to 5 examples. The model must infer the transformation rule and then apply it to a new case.
Unlike MMLU or HumanEval, ARC-AGI 2 rewards neither general knowledge nor syntactic pattern matching. It targets LLM reasoning strictly speaking: rule composition, object manipulation, symmetry, counting. Depending on the Chollet's foundational paper (arXiv:1911.01547), an untrained human exceeds 80% success; the best open-source LLMs plateaued at 5–15% in early 2025.
To understand this benchmark's place in the ecosystem, see our 2025 LLM benchmark guide.
Evaluation protocol: what ARC-AGI 2 really measures
The official protocol imposes two constraints that set ARC-AGI 2 apart from typical benchmarks:
- No data contamination : the 100 private tasks are never published on the internet, preventing any biased pretraining.
- Capped compute budget : each submission is compute-limited to discourage costly brute force.
The official score corresponds to the percentage of tasks solved with two attempts allowed (pass@2). The recommended input format is an ASCII or JSON grid with a color encoding from 0 to 9. The official GitHub repository arc-prize/ARC-AGI-2 provides evaluation scripts, Python harnesses, and reference baselines.
Three families of approaches dominate open-source submissions:
- Direct prompting with structured chain-of-thought
- Test-time training (TTT): lightning-fast fine-tuning on the 2–5 examples before inference
- Program synthesis : Python code generation that the model executes to produce the grid
Test-time training currently delivers the best results, but consumes 2 to 10× more tokens per task. The program synthesis approach was popularized by the MindsAI team in their 2024 submission documented on Arc Prize blog.
Candidate open-source models: VRAM and licenses
Here are the most relevant models in the BestLLMfor catalog for targeting a respectable score on the ARC-AGI 2 benchmark, grouped by hardware profile.
Ultra-high-end (≥ 400 GB Q4 VRAM) :
- DeepSeek V4 Pro 1.6T — 1600B parameters, MIT license, ~960 GB of VRAM in Q4, 1,000,000-token context. MoE architecture suited to multistep reasoning.
- Kimi K2.6 — 1000B, Modified MIT, ~600 GB Q4. Good reported results on planning tasks.
- Ring-1T — 1000B, MIT, ~600 GB Q4, 131,072-token context, designed for long-form reasoning.
- DeepSeek R1 671B — 671B, MIT, ~400 GB Q4. Reasoning model with native chain-of-thought.
- MiMo V2.5 Pro — 1020B, MIT, ~595 GB Q4, 1 M-token context. Interesting for massive few-shot prompts.
Mid-range, accessible to workstations (60–250 GB Q4) :
- Qwen 3 235B-A22B — 235B, Apache 2.0, ~142 GB Q4, context 131,072.
- gpt-oss 120B —117B, Apache 2.0, ~70 GB Q4. Good cost-to-reasoning ratio.
- Nemotron 3 Super 120B — 120B, NVIDIA Open Model License, ~72 GB Q4.
- Llama 4 Scout 109B — 109B, Llama 4 Community, ~65 GB Q4, context extended to up to 10M tokens.
- Mistral Small 4 — 119B, Apache 2.0, ~72 GB Q4, 256,000-token context.
Mainstream profile (≤ 50 GB Q4) :
- Qwen3-Coder-Next 80B-A3B — 80B, Apache 2.0, ~48 GB Q4. Relevant for the program synthesis approach.
- DeepSeek R1 Distill Llama 70B — 70B, ~40 GB Q4. Distilled version of R1’s reasoning.
- Llama 3.3 70B Instruct — 70B, ~40 GB Q4.
Q4 VRAM values are estimates based on GGUF Q4_K_M; for Q8, allow for ×1.9, and for FP16, ×3.8. These figures still need to be confirmed based on the runtime used (llama.cpp, vLLM, SGLang).
Expected performance and tokens/sec
No official public score has yet been consolidated for these models on version 2 of the benchmark—the rankings change every month. For reference, community feedback from theArc Prize Leaderboard 2025 place explicit-reasoning models (R1, Ring-1T) between 8% and 18% pass@2 without TTT, and up to 25–35% with aggressive test-time training. These figures still need to be confirmed for Q4 quantized versions.
In terms of throughput, on a 2× RTX 6000 Ada setup (96 GB total VRAM) with llama.cpp:
- Llama 3.3 70B Q4 : 18–22 tokens/sec (estimated)
- gpt-oss 120B Q4 : 12-16 tokens/sec (estimated)
- Qwen 3 235B-A22B Q4 : 8–12 tokens/sec with partial CPU offload
For detailed reasoning comparisons, see DeepSeek R1 vs Llama 3.3 et the top reasoning models.
Hands-on: reproduce ARC-AGI 2 locally
To evaluate a model from the catalog on the ARC-AGI 2 benchmark, the typical workflow:
- Clone the official repository
arc-prize/ARC-AGI-2and install the Python dependencies. - Load the model through vLLM or llama.cpp in OpenAI-compatible server mode. The vLLM documentation covers MoE models such as Qwen 3 235B-A22B.
- Encode each grid as an ASCII string, then build a prompt with few-shot examples.
- For test-time training, isolate a rank-16–32 LoRA trained on the 2–5 demonstrations for each task (30–90 seconds of overhead per task). The library HuggingFace PEFT covers this use case.
- Evaluate with the script
scoring.pyprovided, which calculates pass@1 and pass@2.
Long-context models such as MiMo V2.5 Pro (1 M tokens) or Llama 4 Maverick 400B (1 M tokens) allow many examples to be injected without truncation, which helps with tasks involving composite rules.
To size your hardware, the BestLLMfor configurator calculates the required VRAM based on the target quantization.
Typical errors on ARC-AGI 2 and approaches
The analysis of failed submissions on the arc-prize/ARC-AGI-2 repository reveals three recurring bug families in open-source LLMs:
- Color hallucination : the model invents a color absent from the 0–9 palette. Mitigation: constrain the output using a typed JSON grammar (logits processor under vLLM).
- Row/column inversion : confusion between dimensions on non-square grids. Mitigation: annotate
H × Lin the prompt. - Overgeneralization : applying a correct rule to the examples but missing the nuance of the test case. Mitigation: force a step
<verification>after chain-of-thought.
For alternative angles of attack, the work by the Greenblatt team (sample-then-filter with Claude Sonnet, see arXiv:2406.07394) shows that a smaller model with massive generation (1000+ sample-and-test Python programs) can outperform a larger model in single-shot mode. This logic transfers well to Qwen3-Coder-Next 80B-A3B, which excels at program synthesis.
Quick comparison with ARC-AGI 1
ARC-AGI 1 (2019) was saturated by late 2024 by hybrid solutions above 55%. ARC-AGI 2 deliberately raises the difficulty again with four changes:
- Composite tasks combining 2 to 4 primitive transformations instead of one.
- Larger grids (up to 30×30 versus 15×15 in v1).
- Removal of statistical biases that enabled brute force through enumeration.
- Private tasks renewed for 2025-2026, including 50 designed by humans who are experts in visual reasoning.
The result: the best submissions plateau at ~30% in 2025, whereas they exceeded 60% on v1. To dig deeper, also see the reasoning vs. math LLM guide and the profile DeepSeek R1 vs. Ring-1T.
FAQ
Q: Which open-source model gets the best ARC-AGI 2 score in late 2025?
As of writing, no official ranking has fixed the open-source podium. Community submissions place DeepSeek R1 671B et Ring-1T leading thanks to their native chain-of-thought. The exact figures remain to be confirmed because the leaderboard is updated monthly by the Arc Prize team.
Q: Do you need a reasoning model or a general-purpose model?
Models with explicit reasoning (R1, Ring-1T, gpt-oss 120B) consistently outperform equivalent general-purpose models with the same parameter count. On ARC-AGI 2, the additional cost in reasoning tokens (often ×3 to ×10) is more than offset by the gain in pass@2. A Llama 3.3 70B standard remains limited compared with a DeepSeek R1 Distill Llama 70B.
Q: What is the minimum VRAM needed to seriously attempt the benchmark?
With 40 GB of VRAM (RTX 6000 Ada or A6000), you can run DeepSeek R1 Distill Llama 70B in Q4 and aim for a respectable score. To reach the open-source frontier, you need to target 250+ GB combined (multi-GPU or aggressive CPU offload) for models such as Qwen 3 235B-A22B.
Q: Does test-time training provide a real gain?
Without TTT, the LLM reasoning relies solely on in-context inference. With a fast LoRA trained on demonstrations for each task, gains of 8 to 15 percentage points are observed. The trade-off is compute costs multiplied by 5 to 10 and increased implementation complexity. For a first test, stick with direct prompting.
Q: Which license applies to commercial use?
Favor permissive licenses. Qwen 3 235B-A22B, gpt-oss 120B, Ring-1T et DeepSeek R1 671B are licensed under Apache 2.0 or MIT, with no commercial-use restrictions. Avoid Command R+ 104B under CC-BY-NC 4.0 for a paid product.
Q: Where can I find the official datasets and scripts?
Everything is centralized on the GitHub repository arc-prize/ARC-AGI-2 and on HuggingFace datasets/arc-prize. The JSON format has been documented and stable since the first iteration in 2019.
Conclusion
Le ARC-AGI 2 benchmark will remain in 2026 one of the few tests resistant to contamination and memorization, making it a valuable tool for anyone who wants to compare the reasoning of open-source LLMs objectively. Before running an evaluation, size your GPUs precisely with the BestLLMfor configurator or browse the full catalog to identify the model best suited to your VRAM budget and target license.