Alibaba has released the Qwen2.5 technical report, introducing a comprehensive series of large language models with significant improvements in both pre-training and post-training stages. The update scales high-quality pre-training datasets from 7 trillion to 18 trillion tokens and implements intricate supervised finetuning with over 1 million samples alongside multistage reinforcement learning.
- Open-weight offerings include base and instruction-tuned models, with quantized versions available.
- Proprietary hosted solutions feature two mixture-of-experts variants: Qwen2.5-Turbo and Qwen2.5-Plus.
- The flagship Qwen2.5-72B-Instruct outperforms many open and proprietary models, showing competitive performance to Llama-3-405B-Instruct.
- Qwen2.5-Turbo and Qwen2.5-Plus offer superior cost-effectiveness while performing competitively against GPT-4o-mini and GPT-4o respectively.
These enhancements provide a strong foundation for common sense, expert knowledge, and reasoning capabilities, while post-training techniques notably improve long text generation, structural data analysis, and instruction following.