xAI commissioned Epoch AI to evaluate Grok 4’s mathematical abilities, characterizing its strengths and weaknesses relative to other models through qualitative investigation and benchmark data.

  • Grok 4 is state-of-the-art at solving medium-hard high school math competitions by grinding out solutions, scoring 88% on benchmarks compared to the previous 85% tie between Gemini 2.5 Pro and o4-mini.
  • The model is near state-of-the-art for proof-based problems from challenging high school competitions but shows significant headroom for improvement in general proofs.
  • Professional mathematicians identified Grok 4 as potentially the best available model for mathematical literature search.
  • The evaluation notes that Grok 4, like other LLMs, favors low-level calculation over spatial intuition or creativity and tends to catch some but not all of its own mistakes.
  • Epoch AI criticized xAI's release blog for misleading performance charts, specifically noting the inappropriate inclusion of Python tools in competitions meant to be taken by hand and a lack of transparency in internal USAMO grading.

This analysis helps identify signs of novel capabilities before they appear in headline numbers and suggests additional benchmarks useful for future model development.