Independent benchmark results released by Epoch AI show that OpenAI’s publicly launched o3 model scored approximately 10% on the FrontierMath dataset, falling significantly short of the over 25% score the company initially highlighted during its December unveiling.

  • Epoch AI evaluated the released o3 model using an updated version of the FrontierMath benchmark and found a score of around 10%.
  • OpenAI’s initial claim of over 25% was achieved using aggressive test-time compute settings on a more powerful internal scaffold than the public release.
  • The ARC Prize Foundation confirmed that the public o3 is tuned for chat and product use, with smaller compute tiers than the preview version they tested.
  • OpenAI technical staff stated the production model is optimized for cost-efficiency and speed rather than maximum benchmark performance.

The discrepancy highlights the importance of verifying third-party benchmarks and clarifies that the released model prioritizes real-world utility over peak academic scores.