OpenAI provided funding for the independent math benchmark FrontierMath while being bound by a non-disclosure agreement that prevented Epoch AI from disclosing this relationship until after OpenAI announced its o3 model's record-breaking performance. The benchmark, developed by over 60 mathematicians, tests complex reasoning skills, and OpenAI's o3 achieved a 25.2 percent success rate on it.
- FrontierMath was introduced in November 2024 to evaluate advanced mathematical problem-solving capabilities.
- Epoch AI signed an agreement preventing disclosure of OpenAI's financial support until the o3 announcement on December 20.
- The benchmark problems were created by a team of over 60 leading mathematicians who were unaware of the funding source.
- OpenAI gained access to "much but not all" FrontierMath data, though a verbal agreement prohibited using the materials for model training.
- Epoch AI acknowledges the lack of transparency as a mistake and promises clearer information on funding sources in future collaborations.
The situation highlights the complexity of AI benchmarking and the importance of transparency, especially since mathematical reasoning is a major weakness for language models.