xAI has announced Grok-1.5, a new model featuring enhanced reasoning capabilities and a context window of up to 128,000 tokens. The model is set to become available to early testers and existing users on the 𝕏 platform in the coming days.

  • Achieved 50.6% on the MATH benchmark and 90% on GSM8K for math tasks.
  • Scored 74.1% on the HumanEval benchmark for code generation.
  • Supports a context length of 128K tokens, increasing memory capacity by 16 times.
  • Demonstrated perfect retrieval results in Needle In A Haystack evaluations.
  • Built on a custom distributed training framework using JAX, Rust, and Kubernetes.

The release aims to provide users with advanced problem-solving abilities and the capacity to process substantially longer documents while maintaining instruction-following accuracy.