Alibaba Cloud's proprietary large language model, Qwen2.5-Max, has achieved the #7 overall ranking on Chatbot Arena, matching other top proprietary LLMs with exceptional capabilities in technical domains.

  • It holds the #1 position in math and coding benchmarks.
  • The model ranks #2 in hard prompts, which involve complex tasks.
  • As a Mixture of Experts (MoE) model, it was pretrained on over 20 trillion tokens and refined via Supervised Fine-Tuning and Reinforcement Learning from Human Feedback.
  • It secured leading scores in major benchmarks including MMLU-Pro, LiveCodeBench, LiveBench, and Arena-Hard.

Developers can access Qwen2.5-Max through Alibaba Cloud's Model Studio or the Qwen Chat platform.