OpenAI's new o3 system, trained on the ARC-AGI-1 Public Training set, has scored 75.7% on the Semi-Private Evaluation set within a $10k compute limit and 87.5% with high-compute configuration.
- The high-efficiency score qualifies as 1st place on the public leaderboard under ARC-AGI-Pub rules.
- Testing was conducted at two compute levels: 6 samples for high-efficiency and 1024 samples (172x compute) for low-efficiency.
- OpenAI confirmed the tested version differs from the officially released o3, with updated results to follow.
ARC Prize views this as a significant step-function increase in AI capabilities, demonstrating novel task adaptation abilities not seen in previous GPT-family models.