The ARC Prize Foundation has published its first public analysis of OpenAI’s o3 and o4-mini models on the ARC-AGI reasoning benchmarks. The evaluation reveals that while these frontier models show significant progress, they have not yet solved the difficult ARC-AGI-2 set.
- o3-medium achieved 53% on the ARC-AGI-1 Semi Private Eval set, while o4-mini-medium reached 41% with high efficiency.
- Both o3 and o4-mini scored below 3% on the more challenging ARC-AGI-2 benchmark.
- High reasoning settings resulted in incomplete coverage due to timeouts or missing outputs, making those results unsuitable for leaderboard reporting.
- The production o3 model differs from the earlier o3-preview by integrating visual inputs and using different training data constraints.
The findings highlight that ARC-AGI-1 remains a sensitive tool for measuring current model capabilities, while ARC-AGI-2 serves as a next-generation benchmark for future models closing the gap between human and AI reasoning.