OpenAI announced its new o3 model, which achieved a 25% score on the Frontier Math benchmark, a significant leap from the previous 2% average for generative AIs. This result highlights rapid improvements in AI mathematical reasoning capabilities.
- Frontier Math is a private benchmark of challenging math questions designed to test human-like abilities without relying on memorized data.
- The o3 model uses iterative chains of thought, allowing it to evaluate and correct its own reasoning steps before answering.
- This process requires substantial computation, with some individual questions costing over $1000 to answer.
- Unlike specialized models like Google DeepMind's AlphaProof, o3 is a general AI that exhibits advanced mathematical skills.
The article suggests that if costs decrease and scores improve, AI mathematicians could transform formal methods by automating verification processes currently limited by the scarcity of human experts.