OpenAI has announced o3, its next-generation reasoning model, which sets new records in reasoning, coding, and mathematical problem-solving. The model achieves a score of 75.7% on the semi-private ARC-AGI eval in low-compute mode and 87.5% in high-compute mode.

  • o3 scores 75.7% on ARC-AGI (low-compute) and 87.5% (high-compute), surpassing the human-level threshold of 85%.
  • It achieves a 96.7% score on the 2024 American Mathematical Olympiad (AIME) and 25.2% on EpochAI’s Frontier Math Benchmark.
  • In coding, o3 scores 71.7 on SWE-Bench Verified (22.8 points higher than o1) and an Elo rating of 2,727 on Codeforces.
  • The model has a context window of 128K tokens and is noted for being significantly more expensive to run than previous iterations.