ARC-AGI-2 Test Challenges AI Models with New Benchmarks

Mar 25, 2025
The Arc Prize Foundation has introduced ARC-AGI-2, a new test designed to evaluate AI models' general intelligence, revealing significant challenges for current AI systems.

The Arc Prize Foundation has launched ARC-AGI-2, a new benchmark test aimed at evaluating the general intelligence of AI models. Announced on their website, the test presents a series of puzzle-like problems that require AI to identify visual patterns and generate correct answers, challenging models to adapt to new problems they haven't encountered before.

Current AI models, including OpenAI's o1-pro and DeepSeek's R1, have scored between 1% and 1.3% on the ARC-AGI-2 test, while non-reasoning models like GPT-4.5 and Claude 3.7 Sonnet scored around 1%. In contrast, human participants averaged a 60% success rate, highlighting the difficulty AI systems face with this new benchmark.

ARC-AGI-2 aims to address the limitations of its predecessor, ARC-AGI-1, by introducing a new metric of efficiency and requiring models to interpret patterns dynamically rather than relying on memorization. The test is designed to measure not only the capability of AI systems to solve tasks but also the efficiency and cost-effectiveness of their solutions.

The Arc Prize Foundation has also announced the ARC Prize 2025 contest, encouraging developers to achieve 85% accuracy on the ARC-AGI-2 test while maintaining a cost of $0.42 per task. This initiative aims to drive open-source progress in developing highly efficient, general AI systems.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Daily AI Brief

Daily report covering major AI developments and industry news, with both top stories and complete market updates

Market report

2025 Generative AI in Professional Services Report

Thomson Reuters

This report by Thomson Reuters explores the integration and impact of generative AI technologies, such as ChatGPT and Microsoft Copilot, within the professional services sector. It highlights the growing adoption of GenAI tools across industries like legal, tax, accounting, and government, and discusses the challenges and opportunities these technologies present. The report also examines professionals' perceptions of GenAI and the need for strategic integration to maximize its value.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.