OpenAI claims its o3-mini model with high reasoning and Python tool access achieves a 32% score on the FrontierMath benchmark. However, Epoch's official independent evaluation reports a significantly lower score of only 11%.

  • Epoch built the benchmark and has better incentives for accurate scoring.
  • OpenAI reported a 28% score on the hardest problem tier, which is suspiciously close to its overall score.
  • Epoch has published extensive information about its testing infrastructure and data, whereas OpenAI has published little.
  • The discrepancy may stem from OpenAI using a more powerful internal scaffold, more test-time compute, or a different subset of problems (180 vs 290).

The difference in results highlights the importance of transparent evaluation methodologies and independent verification in benchmarking large language models.