Independent benchmark results released by Epoch AI show that OpenAI’s publicly launched o3 model scored approximately 10% on the FrontierMath dataset, falling significantly short of the over 25% score the company initially highlighted during its December unveiling.
- Epoch AI evaluated the released o3 model using an updated version of the FrontierMath benchmark and found a score of around 10%.
- OpenAI’s initial claim of over 25% was achieved using aggressive test-time compute settings on a more powerful internal scaffold than the public release.
- The ARC Prize Foundation confirmed that the public o3 is tuned for chat and product use, with smaller compute tiers than the preview version they tested.
- OpenAI technical staff stated the production model is optimized for cost-efficiency and speed rather than maximum benchmark performance.
The discrepancy highlights the importance of verifying third-party benchmarks and clarifies that the released model prioritizes real-world utility over peak academic scores.