OpenAI has previewed its new o3 reasoning language model, signaling rapid progress in AI capabilities ahead of a public release in late January 2025. The announcement highlights significant performance leaps over previous state-of-the-art models like Gemini 1.5 Pro and Claude 3.5 Sonnet.
- o3 is the first model to surpass the 85% threshold on the ARC AGI prize public set, a benchmark designed to measure human-like fluid intelligence.
- It achieves a 25% score on the Frontier Math benchmark, a substantial increase from the previous state-of-the-art of 2%.
- The model reaches International Grandmaster level performance on Codeforces and sets a new state-of-the-art on SWE-Bench-Verified with a 71.7% score.
- OpenAI also released research on deliberative alignment, demonstrating how o1-class models can enhance safety and alignment studies.
These results suggest that reasoning models are beginning to deliver value beyond verifiable domains like mathematics and coding, potentially accelerating progress across the AI research ecosystem.