A review of frontier AI reasoning systems from labs including OpenAI, DeepMind, xAI, DeepSeek, and Anthropic reveals that while test-time compute scaling improves accuracy, it significantly reduces efficiency. The analysis identifies a Pareto frontier for ARC-AGI-1 scores but notes complete failure on the more challenging ARC-AGI-2 benchmark, suggesting common limitations across all major models.
- Modern reasoning systems rely on long-running inference, knowledge recomposition, and parallel CoT sampling to extend "time to think."
- OpenAI's o3-medium/high offers highest accuracy at high cost, while Google's Gemini 2.5 Flash provides a balance of accuracy and cost.
- xAI's Grok 3 mini-low offers modest accuracy improvements over pure LLMs with lower resource requirements.
- ARC-AGI-2 remains unsolved by all tested systems, indicating that current scaling strategies are insufficient for breakthrough AGI capability.
The findings suggest there is no single universal winner and that AI reasoning systems lack strong product-market fit compared to pure LLMs due to API inconsistencies and adoption friction.