OpenAI's new o3 system, trained on the ARC-AGI-1 Public Training set, has scored 75.7% on the Semi-Private Evaluation set within a $10k compute limit and 87.5% with high-compute configuration.

  • The high-efficiency score qualifies as 1st place on the public leaderboard under ARC-AGI-Pub rules.
  • Testing was conducted at two compute levels: 6 samples for high-efficiency and 1024 samples (172x compute) for low-efficiency.
  • OpenAI confirmed the tested version differs from the officially released o3, with updated results to follow.

ARC Prize views this as a significant step-function increase in AI capabilities, demonstrating novel task adaptation abilities not seen in previous GPT-family models.