xAI has released an early preview of Grok-2 and its smaller variant, Grok-2 mini, marking a significant upgrade from the previous Grok-1.5 model. These new models are currently available in beta on 𝕏 for Premium and Premium+ users, with enterprise API access planned for later this month.

  • Grok-2 demonstrates frontier capabilities in chat, coding, and reasoning, outperforming Claude 3.5 Sonnet and GPT-4-Turbo on the LMSYS Chatbot Arena leaderboard.
  • Both models show significant improvements over Grok-1.5 in following instructions, providing factual information, and tool use.
  • Benchmarks indicate competitive performance with other frontier models in graduate-level science (GPQA), general knowledge (MMLU, MMLU-Pro), and math (MATH).
  • Grok-2 achieves state-of-the-art results in visual math reasoning (MathVista) and document-based question answering (DocVQA).
  • The enterprise API will support multi-region inference deployments with enhanced security features like mandatory multi-factor authentication.

The release positions xAI at the forefront of AI development by advancing core reasoning capabilities through a new compute cluster and expanded multimodal understanding.