Alibaba's Qwen team has released Qwen3.8-Omni-Flash, its first omni-modal model designed around agentic capabilities for audio-video understanding and tool use. The model accepts text, images, audio, and video inputs to return text outputs, featuring a 1M-token context window and native function calling.
- Built on the Qwen3.8-Flash-Next architecture, it offers a 1M-token context with 991K max input and 131K max output tokens.
- It employs an agentic perception workflow for long videos, raising OmniVideoBench accuracy from 63.4 to 67.8 while reducing token usage by 45.7%.
- The model shows significant performance gains over Qwen3.5-Omni-Plus, with average scores improving more than 25% across 29 evaluations.
- It is available via API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio, with pricing at $0.15 per 1M input tokens and $0.47 per 1M output tokens.
- The team also open-sourced Qwen-MM-Plugins under Apache-2.0 to support multimodal agent harnesses.
The release provides a hosted solution for complex agentic workflows involving long-form audio and video analysis, with reported cost reductions of over 93% for audio-visual inputs compared to previous models.