A comparative analysis evaluates the capabilities of Claude 3.7 Sonnet, Claude 3.5 Sonnet, OpenAI o3-mini, DeepSeek R1, and Grok 3 Beta across math, coding, and reasoning benchmarks.

  • On GPQA Diamond, Grok 3 Beta reaches 84.6% with extended reasoning, while Claude 3.7 Sonnet achieves 84.8% in the same mode.
  • In AIME math competitions, Grok 3 Beta leads with 93.3%, whereas Claude 3.7 Sonnet scores 61.3% but excels on Math500 at 96.2%.
  • OpenAI o3-mini demonstrates strong performance on AIME (up to 83.3%) and GPQA Diamond (78% in higher-compute mode) with a focus on cost-effectiveness.
  • DeepSeek R1 performs competitively on AIME (79.8%) but shows variability on graduate-level reasoning tasks.

The study aims to help users make informed decisions by highlighting trade-offs in cost, context window size, and reliability across these leading models.