Alibaba has released the Qwen2.5 technical report, introducing a comprehensive series of large language models with significant improvements in both pre-training and post-training stages. The update scales high-quality pre-training datasets from 7 trillion to 18 trillion tokens and implements intricate supervised finetuning with over 1 million samples alongside multistage reinforcement learning.

  • Open-weight offerings include base and instruction-tuned models, with quantized versions available.
  • Proprietary hosted solutions feature two mixture-of-experts variants: Qwen2.5-Turbo and Qwen2.5-Plus.
  • The flagship Qwen2.5-72B-Instruct outperforms many open and proprietary models, showing competitive performance to Llama-3-405B-Instruct.
  • Qwen2.5-Turbo and Qwen2.5-Plus offer superior cost-effectiveness while performing competitively against GPT-4o-mini and GPT-4o respectively.

These enhancements provide a strong foundation for common sense, expert knowledge, and reasoning capabilities, while post-training techniques notably improve long text generation, structural data analysis, and instruction following.