The Qwen Team at Alibaba Group has published a technical report detailing Qwen2.5-VL, the latest flagship model in their vision-language series. The report outlines significant advancements in visual recognition, object localization, document parsing, and long-video comprehension.
- The model introduces dynamic resolution processing and absolute time encoding to handle images of varying sizes and videos up to hours long with second-level event localization.
- A native dynamic-resolution Vision Transformer (ViT) with Window Attention reduces computational overhead while maintaining native resolution.
- Pre-training corpus was scaled from 1.2 trillion tokens to 4.1 trillion tokens, enhancing performance across domains without task-specific fine-tuning.
- The flagship Qwen2.5-VL-72B model matches state-of-the-art models like GPT-4o and Claude 3.5 Sonnet, while smaller 7B and 3B variants outperform competitors in resource-constrained environments.
The report emphasizes the model's ability to serve as an interactive visual agent for reasoning and task execution on computers and mobile devices.