Epoch AI has introduced FrontierMath, a benchmark consisting of hundreds of original, expert-crafted mathematics problems designed to assess advanced reasoning capabilities in AI systems. Developed in collaboration with over 60 mathematicians, including Fields medalists, the problems span modern research areas like number theory and algebraic geometry, typically requiring hours or days for experts to solve.

  • The benchmark features problems that are "guessproof" and automatically verifiable through computation.
  • Evaluation of six leading models, including Claude 3.5 Sonnet and GPT-4o, showed none could solve more than 2% of the problems despite extensive support frameworks.
  • This performance contrasts sharply with benchmarks like GSM-8K, where top models achieve over 90% accuracy.
  • Epoch AI plans to conduct regular evaluations, expand the benchmark, and release additional problems in the coming months.

FrontierMath aims to provide a standardized measure of progress in research-level mathematical reasoning as AI systems advance, highlighting a substantial gap between current AI capabilities and expert mathematicians.