OpenAI’s new language model, o3, achieved a 25% score on FrontierMath, a secret dataset of hundreds of difficult mathematics questions curated by Epoch AI. The dataset consists of problems with definitive, computable answers that can be automatically verified, designed to prevent models from simply memorizing public solutions.

  • FrontierMath contains "hundreds" of hard math questions, with fewer than 200 initially released and more added subsequently.
  • The five publicly available sample problems require positive whole number answers, such as 9811 and 367707, and involve complex concepts like p-adic continuity and Weil conjectures.
  • While the author could solve two of the five public questions using their expertise in arithmetic, they noted that a typical smart mathematics undergraduate would likely struggle with even one.
  • Epoch AI's Elliot Glazer suggested that 25% of the dataset consists of "IMO/undergrad style problems," raising questions about the overall difficulty level and representativeness of the public samples.

The achievement is significant because the area of AI for mathematics lacks hard datasets, and creating such benchmarks is difficult and expensive. The author expressed shock at the score, noting that while AI is improving rapidly on Olympiad-style problems, reaching advanced undergraduate or PhD-level innovation was previously considered a distant goal.