The ARC Prize Foundation has published its first public analysis of OpenAI’s o3 and o4-mini models on the ARC-AGI reasoning benchmarks. The evaluation reveals that while these frontier models show significant progress, they have not yet solved the difficult ARC-AGI-2 set.

  • o3-medium achieved 53% on the ARC-AGI-1 Semi Private Eval set, while o4-mini-medium reached 41% with high efficiency.
  • Both o3 and o4-mini scored below 3% on the more challenging ARC-AGI-2 benchmark.
  • High reasoning settings resulted in incomplete coverage due to timeouts or missing outputs, making those results unsuitable for leaderboard reporting.
  • The production o3 model differs from the earlier o3-preview by integrating visual inputs and using different training data constraints.

The findings highlight that ARC-AGI-1 remains a sensitive tool for measuring current model capabilities, while ARC-AGI-2 serves as a next-generation benchmark for future models closing the gap between human and AI reasoning.