OpenAI claims its o3-mini model with high reasoning and Python tool access achieves a 32% score on the FrontierMath benchmark. However, Epoch's official independent evaluation reports a significantly lower score of only 11%.
- Epoch built the benchmark and has better incentives for accurate scoring.
- OpenAI reported a 28% score on the hardest problem tier, which is suspiciously close to its overall score.
- Epoch has published extensive information about its testing infrastructure and data, whereas OpenAI has published little.
- The discrepancy may stem from OpenAI using a more powerful internal scaffold, more test-time compute, or a different subset of problems (180 vs 290).
The difference in results highlights the importance of transparent evaluation methodologies and independent verification in benchmarking large language models.